HBase writes are fast because they never update a file in place. A write goes to the write-ahead log and the in-memory MemStore. When the MemStore fills, it is flushed as a new immutable HFile. Every flush adds a file, and every file is one more place a read may have to look. Compaction is the background process that merges those files back together. It is what keeps reads fast and removes deleted and expired data, and it is also the largest source of disk and network I/O on most busy clusters.

The reasons compaction exists, and the read, write and space amplification trade-off behind it, are covered in the HBase compaction overview. This page is the practical one. It covers how the default policy chooses files, what each configuration property does and what its real default is, how compaction throttling and write blocking interact, which alternative store engines exist, and a worked tuning exercise on a write-heavy table. All defaults quoted here were checked against hbase-default.xml and the compaction source on the HBase 2.6 branch. If you run another release line, check your own defaults, because several changed between 1.x and 2.x.

The compaction path in a RegionServer

Each column family in each region is a Store, and each Store has its own set of HFiles. Compaction happens per Store, so a table with three column families and 200 regions has 600 independent piles of files. That is one reason keeping column families few matters for compaction cost.

After every flush, and periodically from a background chore, the RegionServer asks each Store whether it needs compaction. If the Store has at least the minimum number of eligible files, the selection policy picks a set and the request is queued. Requests whose total size is above the throttle point go to the large pool, and the others go to the small pool. Each pool has one thread by default. The merge then reads the selected files in key order, writes one new HFile, and swaps it in atomically.

There are two kinds of compaction. A minor compaction merges a subset of files, usually the newer and smaller ones. It cannot drop delete markers in general, because an older file outside the selection might still hold the data they mask. A major compaction rewrites all files in the Store into one. It can drop deleted cells, cells past their TTL and versions beyond the column family's limit. Major compactions reclaim space and restore data locality, but they rewrite everything, which is why when they run matters so much.

MemStore flush~128 MB HFileStore (one CF)N HFilesSelection policyratio, min, maxCompaction queuesmall / largeSmall pool1 thread defaultLarge pool1 thread defaultThroughput controller50 to 100 MB/s, scaled by file pressureBlocking checkfiles >= 16: flush waitsWritersstall up to 90 snew filerate limitcountmerged filelarge vs small: total selected size compared with hbase.regionserver.thread.compaction.throttle (2.5 GiB default)
Figure 2. Flushes add files; the policy selects, the queue routes by size, the throughput controller limits I/O, and too many files block flushes.

How files are selected

The default policy in HBase 2.x is ExploringCompactionPolicy. It looks at the eligible files and considers every contiguous run of between hbase.hstore.compaction.min and hbase.hstore.compaction.max files. For each candidate set, every file must pass the ratio test: its size must be no more than hbase.hstore.compaction.ratio times the combined size of the other files in the set. Files smaller than hbase.hstore.compaction.min.size always pass, and files larger than hbase.hstore.compaction.max.size are never selected. Among the sets that pass, the policy prefers the one with the most files, and on a tie the smallest total size. That rewrites the least data for the most reduction in file count.

The ratio is the main control. With a ratio of 1.2 and files of 400, 100, 60, 50 and 40 MB, the 400 MB file fails, because 400 is more than 1.2 times 250 (the sum of the others). The four small files pass and are merged into one 250 MB file. A higher ratio lets large files join merges more often, giving fewer files and better reads at the cost of rewriting more data. A lower ratio does the opposite.

def passes_ratio(files_mb, ratio, min_size_mb=128):
    # Every file must be small enough relative to the rest of the candidate set.
    total = sum(files_mb)
    for f in files_mb:
        if f > min_size_mb and f > ratio * (total - f):
            return False
    return True

def explore(files_mb, min_files=3, max_files=10, ratio=1.2):
    best = None
    for start in range(len(files_mb)):
        for end in range(start + min_files, min(len(files_mb), start + max_files) + 1):
            cand = files_mb[start:end]
            if passes_ratio(cand, ratio):
                key = (-len(cand), sum(cand))     # most files, then least bytes
                if best is None or key < best[0]:
                    best = (key, cand)
    return best and best[1]

print(explore([400, 100, 60, 50, 40]))           # [100, 60, 50, 40]

This is a simplified model. The real policy also excludes files already being compacted, handles bulk-loaded files and applies the off-peak ratio when the current hour is inside the off-peak window. It is accurate enough to reason about what a ratio change will do.

The properties that matter

These are the properties that matter, with the defaults from hbase-default.xml and the source on the 2.6 branch:

PropertyDefaultEffect
hbase.hstore.compaction.min3 (older name: hbase.hstore.compactionThreshold)Minimum files per minor compaction; higher means fewer, bigger merges
hbase.hstore.compaction.max10Maximum files per minor compaction
hbase.hstore.compaction.ratio1.2Ratio test during normal hours
hbase.hstore.compaction.ratio.offpeak5.0Ratio during off-peak hours, which is far more aggressive
hbase.offpeak.start.hour / end.hour-1 (disabled)Off-peak window, hours 0 to 23
hbase.hstore.compaction.min.size128 MBFiles below this always pass the ratio test
hbase.hstore.compaction.max.sizeLong.MAX_VALUEFiles above this are never minor-compacted
hbase.hregion.majorcompaction604800000 ms (7 days)Interval for periodic major compaction; 0 disables it
hbase.hregion.majorcompaction.jitter0.5Spreads major compactions across the interval
hbase.hstore.blockingStoreFiles16File count at which flushes are delayed
hbase.hstore.blockingWaitTime90000 msMaximum time a flush is delayed
hbase.regionserver.thread.compaction.small / large1 / 1Compaction threads per pool
hbase.regionserver.thread.compaction.throttle2684354560 (2.5 GiB)Total size above which a request goes to the large pool
hbase.hstore.compaction.throughput.lower.bound52428800 (50 MB/s)Throughput limit when there is no file pressure
hbase.hstore.compaction.throughput.higher.bound104857600 (100 MB/s)Throughput limit as pressure approaches blocking

Most of these can also be set per table or per column family, which is often better than changing them for the whole cluster:

# HBase shell: override compaction behaviour for one hot table only
alter 'events', CONFIGURATION => {
  'hbase.hstore.compaction.min'   => '5',
  'hbase.hstore.compaction.ratio' => '1.4',
  'hbase.hstore.blockingStoreFiles' => '30'
}

Throttling and blocking

Compaction competes with reads and flushes for disk and network. In 2.x the default throughput controller is PressureAwareCompactionThroughputController. It sets a per-RegionServer compaction rate limit between the lower and higher bounds based on file pressure, which is roughly how close the Stores are to the blocking file count. With few files, compactions are held near 50 MB/s so they do not disturb serving. As files accumulate, the limit rises towards 100 MB/s. If pressure goes beyond the blocking point, the limit is removed so the server can catch up. During the off-peak window the limit is controlled by hbase.hstore.compaction.throughput.offpeak, which defaults to unlimited.

The blocking mechanism is the safety valve, and it is where users feel compaction. When a Store reaches hbase.hstore.blockingStoreFiles files, a flush for that region is delayed until compaction reduces the count or hbase.hstore.blockingWaitTime expires. While the flush waits, the MemStore keeps filling. When it reaches its blocking size, writes to the region are rejected with a busy exception and clients retry. To the application this looks like random write latency spikes of tens of seconds. In the RegionServer log you will see messages about too many store files and delayed flushes.

Raising blockingStoreFiles is the common first response, and it is sometimes right, because 16 is conservative for large write-heavy regions. But it only moves the wall: reads get slower as file counts grow, and if compaction cannot keep up at all, a higher limit just delays the stall. The real fix is matching compaction capacity to the flush rate, as the worked example below shows. For the flush side of this equation, see MemStore flushing.

Taking control of major compactions

Periodic major compaction with jitter means each Store rewrites all its data about once a week at a time HBase chooses. On a cluster with tens of terabytes, that can mean several large rewrites at peak hours. Many operators therefore set hbase.hregion.majorcompaction to 0 and trigger major compactions themselves:

# Disable time-based major compaction (hbase-site.xml)
#   hbase.hregion.majorcompaction = 0
# Then schedule your own, for example from cron in a quiet window:
echo "major_compact 'events'" | hbase shell -n
echo "compaction_state 'events'" | hbase shell -n     # NONE / MINOR / MAJOR / MAJOR_AND_MINOR

# During an incident, pause compactions on one RegionServer, then resume
echo "compaction_switch false, ['rs1.example.com,16020,1696300000000']" | hbase shell -n
echo "compaction_switch true,  ['rs1.example.com,16020,1696300000000']" | hbase shell -n

If you disable periodic majors, you must actually run them. Otherwise delete markers and expired cells accumulate, space grows and reads slow down because scans must skip ever more tombstones. Rolling major compactions table by table or region by region, a few at a time, keeps the I/O bounded. Remember that minor compactions can be promoted to major ones when they happen to select every file in a Store, so disabling periodic majors does not stop all full rewrites.

Date-tiered and stripe store engines

The default store engine treats all files alike. Two alternatives change the selection model for specific data shapes:

  • Date-tiered compaction (hbase.hstore.engine.class set to org.apache.hadoop.hbase.regionserver.DateTieredStoreEngine) groups files into time windows and compacts only within a window. Windows grow exponentially with age: hbase.hstore.compaction.date.tiered.base.window.millis defaults to 6 hours and windows.per.tier to 4. This suits time-series data that is written once, read mostly by recent time range and expired by TTL, because old windows stop being rewritten and whole files can expire together. It suits data with out-of-order timestamps or frequent updates to old rows poorly.
  • Stripe compaction (org.apache.hadoop.hbase.regionserver.StripeStoreEngine) splits a large region into key-range stripes, each compacted on its own. It reduces the size of individual compactions in very large regions and can help when regions must be large but writes are spread across the key range.

Both are per-table settings applied with alter. Test them on a copy of real data, because the benefit depends entirely on the key and timestamp distribution.

Worked example: a write-heavy table that stalls

Consider an events table on 12 RegionServers ingesting 300 MB/s of raw data across 1,200 regions, one column family and a 128 MB flush size. Users report write stalls of 20 to 60 seconds every few minutes. The metrics show compactionQueueLength on several servers stuck in the dozens, and many Stores at 16 files.

Work through the arithmetic. Each server absorbs about 25 MB/s of flushes. With minor compactions of 3 files, each byte is rewritten several times as small files merge into larger ones; a write amplification of 4 to 6 is typical. So each server needs 100 to 150 MB/s of compaction throughput, while the controller allows 50 to 100 MB/s and there are only two compaction threads. Compaction simply cannot keep up, so files pile up until flushes block.

The changes, in order of safety:

  1. Raise hbase.regionserver.thread.compaction.small to 3 and large to 2. One thread per pool is far too few for 100 regions per server.
  2. Raise the throughput bounds to 100 and 200 MB/s after checking disk headroom.
  3. Raise hbase.hstore.compaction.min to 5 for this table, so each merge reduces the file count more per byte written.
  4. Raise blockingStoreFiles to 30 for this table to absorb bursts, now that compaction can drain them.
  5. Disable periodic majors, and run them region by region at night.

After a day, queue length settles in single digits, average files per Store is around 6, and write stalls stop. Read p99 rises slightly because of the extra files during the day, which is the trade being made consciously. If reads matter more than writes for a table, take the opposite path: a lower minimum, a higher ratio and more frequent majors.

What to watch

Tune from metrics, never from intuition. The RegionServer exposes compaction queue length, store file count per region, compacted bytes and compaction time through JMX and the web UI, and the HBase metrics guide covers how to collect them. Watch queue length (it should drain to near zero between bursts), maximum store files per Store compared with the blocking limit, flush delays and the busy exceptions seen by clients, and read p99 alongside file count. A sustained rise in files per Store with a flat write rate means compaction has fallen behind or a large file is failing the ratio test repeatedly.

Failure modes

  • Compaction storm after recovery. When a RegionServer restarts, many Stores may request compaction at once. Throughput limits protect serving, but queue drain can take hours.
  • Majors at peak. Default periodic majors with jitter still start whenever the timer fires. Disable them or move them deliberately.
  • Raising blocking only. Hides the backlog until reads degrade and the stall returns at a higher file count.
  • Too many column families. Each family flushes and compacts separately, multiplying small files.
  • Huge regions. A major compaction of a 50 GB region is a 50 GB rewrite. Region sizing and compaction tuning are the same problem; see region count and sizing.
  • Tombstone build-up. Heavy deletes with no major compaction make scans slower even though space looks stable.

What to do next

  1. Record your cluster's actual values for every property in the table above, from the RegionServer configuration page rather than from memory.
  2. Graph compaction queue length and maximum store files per Store for a week, and mark where blocking happened.
  3. Estimate flush rate times write amplification per server and compare it with the throughput bounds and thread count.
  4. Move tuning for your hottest table into a per-table CONFIGURATION override, change one property at a time, and measure.
  5. Decide whether to disable periodic majors, and if you do, script and monitor your own schedule.
  6. Evaluate date-tiered compaction on a staging copy of any TTL-bound time-series table.
Key takeaway: Compaction tuning is capacity planning: compaction throughput and threads must cover the flush rate times write amplification, or files pile up until flushes block. Read the ratio test to predict selection, tune per table, raise blocking limits only after compaction can drain, schedule major compactions yourself, and verify every change against queue length and file-count metrics.