HBase never updates a file in place. Every write lands in the write-ahead log and an in-memory MemStore, and when the MemStore fills it is flushed as a new immutable HFile. Left alone, a busy column family would accumulate hundreds of HFiles, and every read would have to consult all of them. Compaction is the background process that merges those files back together, and HBase has two kinds of it that behave very differently: minor compactions, which merge a few files cheaply and leave deletes in place, and major compactions, which rewrite every file in a store and are the only point at which deleted and expired data is guaranteed to leave the disk.

This article is about that split. It explains, from first principles, what each kind reads, writes and drops; how the default ExploringCompactionPolicy chooses files; when a minor compaction silently becomes a major one; how the weekly major schedule and its jitter work; and what happens to writers when compaction falls behind. It finishes with a worked example, a list of failure modes and a checklist. For the wider LSM write path see HBase compaction: taming the LSM write path.

The model: stores, HFiles and tombstones

A table is split into regions, each region holds one store per column family, and each store is a MemStore plus a set of HFiles. Every cell carries a row, family, qualifier, timestamp and type, and HFiles are sorted by that key. A read builds a merging scanner over the MemStore and every HFile that might contain the row, skipping files using Bloom filters and timestamp ranges where it can. The cost of a read is therefore roughly proportional to the number of files it cannot rule out.

Deletes are writes too. A Delete inserts a tombstone (a cell of type Delete, DeleteColumn or DeleteFamily) that masks older cells at read time. The masked data still sits in older HFiles. It can only be physically removed when the tombstone and every cell it masks are rewritten in the same pass, because a tombstone dropped too early would let an older cell in an untouched file reappear. That single constraint explains most of the difference between the two kinds of compaction.

MemStoresorted, in memoryflushHFile 1 (old)HFile 2HFile 3HFile 4 (new)minor: 2-4Merged HFiledeletes keptMajor compactionall files -> one; deletes and expired data droppedSingle HFilelocality restoredRead path cost= files consultedper Get/ScanBloom filters andtime-range skip help,but fewer files isthe real fix
Flushes create small HFiles. Minor compactions merge a window of them and keep tombstones; a major compaction rewrites the whole store into one file and drops deleted and expired data.

Minor compactions

A minor compaction picks a subset of a store's files, merges them with a single sequential pass, and writes one new HFile. Because the selected files are not the whole store, there may be older files outside the selection that hold cells masked by a tombstone inside it. HBase therefore runs a minor compaction with a scan type that retains delete markers. The Apache HBase reference guide is blunt about it: do not rely on minor compactions to remove deleted or expired data; only a major compaction guarantees it.

What minor compactions buy you is read amplification control at low cost. Merging four 128 MB flush files reads and writes about 512 MB, rather than rewriting a 20 GB store. They run constantly on a busy cluster and are triggered after flushes, when a store has at least hbase.hstore.compaction.min eligible files (default 3; the older name hbase.hstore.compactionThreshold still works). A single minor compaction takes at most hbase.hstore.compaction.max files (default 10).

Major compactions

A major compaction selects every file in a store and writes one output file. Because nothing older can exist outside the selection, it is safe to drop tombstones together with the cells they mask, to drop versions beyond the family's VERSIONS setting, and to drop cells older than the family's TTL (keeping at least MIN_VERSIONS if set). If a family has KEEP_DELETED_CELLS => TRUE, deleted cells survive so that time-range queries can still see them, until TTL or the version limit removes them.

A major compaction also restores HDFS data locality. The output is written by the RegionServer that hosts the region, so the first replica of every block lands on the local DataNode. After a rolling restart or a balancer run, a major compaction is the most thorough (and most expensive) way to make reads local again; see HBase region data locality.

The price is write amplification. A region with 30 GB in one family rewrites 30 GB of data, through the HDFS write pipeline and its three replicas, for every major compaction. On a cluster with a few thousand regions that is a large share of daily disk and network capacity.

How files are selected

Since HBase 0.96 the default policy is ExploringCompactionPolicy. It treats a store's files as a list ordered from oldest to newest and examines every contiguous window whose length lies between the minimum and maximum file counts. A window is acceptable only if no file in it is too large compared with the others, the ratio test:

# pseudocode for ExploringCompactionPolicy.selectFiles (simplified)
candidates = [f for f in store_files                       # oldest -> newest
              if f.size <= COMPACTION_MAX_SIZE and not f.being_compacted]
ratio = RATIO_OFFPEAK if in_offpeak_hours() else RATIO     # 5.0 vs 1.2
best = None
for start in range(len(candidates)):
    for end in range(start + MIN_FILES, min(start + MAX_FILES, len(candidates)) + 1):
        window = candidates[start:end]
        ok = all(f.size <= MIN_SIZE or                    # small files always pass
                 f.size <= ratio * (total(window) - f.size)
                 for f in window)
        if not ok:
            continue
        # prefer more files; on a tie, prefer the smaller total size
        if best is None or len(window) > len(best) or \
           (len(window) == len(best) and total(window) < total(best)):
            best = window
if best is None and store_is_stuck():                      # too many files
    best = smallest_window_of(MIN_FILES, candidates)
return best

Read the ratio test as: a file joins a compaction only if it is no more than 1.2 times the combined size of the other files in the window. A freshly major-compacted 20 GB file fails that test against three 128 MB flushes, so it is left alone, which is exactly what you want. Files under hbase.hstore.compaction.min.size (default 128 MB) pass automatically, and files over hbase.hstore.compaction.max.size (default Long.MAX_VALUE, effectively unlimited) are never selected for a minor compaction.

During off-peak hours, defined by hbase.offpeak.start.hour and hbase.offpeak.end.hour (both -1, disabled, by default), the ratio becomes hbase.hstore.compaction.ratio.offpeak (default 5.0), letting bigger files join and producing fewer, larger merges while traffic is low.

When a compaction becomes major

Two paths turn a compaction into a major one. The first is promotion: if the policy's selection happens to include every file in the store, the compaction is run as a major, tombstones and all. Small stores with few files are promoted often, which is harmless. The second is time: each store records when it was last major-compacted, and once hbase.hregion.majorcompaction has elapsed (default 604800000 ms, seven days) the next compaction check requests a major one. To stop every store on a cluster becoming due in the same hour, the interval is spread by hbase.hregion.majorcompaction.jitter (default 0.50), so a store's actual period falls somewhere between about 3.5 and 10.5 days.

Setting hbase.hregion.majorcompaction to 0 disables time-based majors. Many operators do that and trigger majors themselves at a quiet hour, table by table, so the cost is predictable:

# hbase shell
major_compact 'events'                 # every region of a table
major_compact 'events', 'd'            # one column family
compact 'events'                       # request a minor compaction
compaction_state 'events'              # NONE, MINOR, MAJOR or MAJOR_AND_MINOR
compaction_switch false                # pause compactions on all RegionServers

Note that a major request is still queued and throttled like any other compaction; it does not run instantly. Turning compactions off with compaction_switch false also interrupts compactions already in progress, as does closing or moving the region.

Back-pressure, thread pools and throttling

Compaction is a background job, but it is not optional. If flushes outpace compactions, a store's file count grows, and once it exceeds hbase.hstore.blockingStoreFiles (default 16 in current releases) the region blocks further flushes. The MemStore keeps filling, and when it reaches the blocking multiple of its flush size, writes to that region stall with RegionTooBusyException. The block lifts when compaction brings the file count down, or after hbase.hstore.blockingWaitTime (default 90000 ms) elapses.

Each RegionServer runs compactions in two thread pools. Selections whose total size exceeds hbase.regionserver.thread.compaction.throttle go to the large pool; the rest go to the small pool. The default threshold, 2684354560 bytes, is 2 x compaction.max x the 128 MB flush size, i.e. 2.5 GiB. Both pools default to one thread, so one huge major compaction cannot starve all the small merges that keep file counts in check. Raising the small pool to two or three threads is a common fix for write-heavy clusters.

HBase 2.x also throttles compaction I/O. The pressure-aware throughput controller targets between hbase.hstore.compaction.throughput.lower.bound (52428800 bytes/s, 50 MB/s) and hbase.hstore.compaction.throughput.higher.bound (104857600 bytes/s, 100 MB/s) per RegionServer, moving toward the upper bound as store-file pressure rises. If compactions fall behind on fast disks, the bounds rather than the thread counts are often the limit.

Worked example: sizing one RegionServer

Take a RegionServer hosting 100 regions of an events table with one family, a 7-day TTL and steady writes of 2.5 MB/s, about 216 GB a day. Each region receives about 25 KB/s, so its MemStore never reaches 128 MB within an hour; the hourly periodic flush writes files of roughly 90 MB instead.

  • With compaction.min = 3, each region merges three flush files about every three hours. A merged file passes the ratio test against two others of similar size, so files grow in roughly geometric steps. Minor compaction I/O works out at a small multiple of ingest, perhaps 10-20 MB/s of combined reads and writes, comfortably under the 100 MB/s throughput ceiling.
  • With a 7-day TTL each region settles at about 15 GB of live data, but only if expired cells actually leave the disk, and that is reliably done by major compactions. A weekly major on 100 regions rewrites about 1.5 TB per server; at 100 MB/s that is about four hours of continuous compaction, which is why scheduling it at night matters.
  • If you disable time-based majors and forget to schedule your own, expired cells and tombstones accumulate: storage grows well past the 15 GB the TTL implies and scans slow down as they skip more and more dead cells.
  • Now scale ingest to 40 MB/s on the same server. Minor compaction I/O alone would need well over 100 MB/s, so with default bounds compaction falls behind, file counts climb toward blockingStoreFiles and writes stall. Raise the throughput bounds and thread counts, or add RegionServers, before ingest grows that far.

The plan that follows: disable time-based majors, enable off-peak hours from 01:00 to 05:00, raise the small pool to 2 threads, and run a script that major-compacts a quarter of the tables each night, so every table is rewritten about twice a week.

Failure modes

  • Compaction storm after a restart. If many stores are near their seven-day mark, a cluster-wide restart can make a large share of them due at once. Jitter helps only if majors were not all triggered manually at the same moment last week. Stagger your own schedule.
  • Write stalls. Logs show Waited ... ms on a compaction to clean up 'too many store files' and clients get RegionTooBusyException. Look at the compaction queue length metric and store file counts per region before raising blockingStoreFiles; raising it only defers the stall and makes reads slower.
  • Deleted data that will not go away. Storage does not shrink after a mass delete because minor compactions keep tombstones. Run a major compaction on the affected family, and check KEEP_DELETED_CELLS is not set.
  • Snapshots pin old files. A compaction's input files are archived, not deleted, while a snapshot references them. Old snapshots after a major compaction can double a table's footprint on HDFS.
  • Majors that ruin latency. A major on a large store competes with reads for disk, and the new file starts with a cold block cache. Prefer off-peak scheduling, and consider cache-on-write settings for hot tables.

Trade-offs

ChoiceYou gainYou pay
Lower compaction.minFewer files, faster readsMore write amplification
Higher compaction.ratioLarger merges, fewer filesBig files rewritten more often
Time-based majors onNo forgotten tombstonesUnpredictable heavy I/O
Manual scheduled majorsPredictable windowsYou own the job and its failures
Higher blockingStoreFilesFewer write stallsSlower reads, longer catch-up

For time-series tables where data is only ever appended and expired by TTL, Date Tiered or FIFO compaction policies can avoid most rewriting; they are separate policies with their own constraints and are worth reading about before you tune the default one harder.

What to do next

  1. Record hbase.hregion.majorcompaction, jitter, off-peak hours and thread-pool sizes from your hbase-site.xml; know whether time-based majors are on.
  2. Graph compaction queue length, store files per region and the count of write stalls; alert when any region exceeds about 75 percent of blockingStoreFiles.
  3. If you disable time-based majors, write and schedule the replacement job in the same change, and alert when a table has not been major-compacted in 10 days.
  4. After any bulk delete, run a family-scoped major_compact and confirm the store size falls.
  5. Read the HFile format and the RegionServer internals to understand what each merge reads and writes, and regions and splits for how region size drives major cost.
Key takeaway: Minor compactions keep the file count down cheaply but never remove deleted data; major compactions rewrite whole stores, drop tombstones and expired versions, and restore locality at a high I/O cost. Know which schedule you run, watch store-file counts against the blocking limit, and plan majors at quiet hours.