In most Java services a long garbage-collection pause is a latency problem. In HBase it is an availability problem. A RegionServer that stops for longer than its ZooKeeper session timeout is declared dead, its regions are reassigned to other servers, and when it wakes up it discovers it has been fenced off and aborts. A single bad pause can therefore take a server out of the cluster and trigger a wave of region moves and WAL splitting that hurts every other server too.

This article explains GC tuning for HBase 2.x RegionServers from the allocation pattern upwards: what lives in the heap and for how long, how MSLAB prevents fragmentation, which G1 settings matter and why, when to move data off-heap, how to read a GC log, and how to budget a real server's memory. It ends with a failure table and a checklist. Defaults quoted come from hbase-default.xml on the 2.6 branch; verify them against the version you run.

Advertisement

Why a pause looks like death

A RegionServer holds a session with ZooKeeper and keeps it alive with heartbeats. HBase sets zookeeper.session.timeout to 90,000 ms by default, but that is only a request: the ZooKeeper server caps sessions at its own maxSessionTimeout, which defaults to 20 times its tick time, or 40 seconds with the usual 2-second tick. Unless the ensemble has been configured to allow more, your real timeout is 40 seconds, not 90. This mismatch surprises many operators.

When the JVM stops every thread for a full collection, heartbeats stop too. If the pause outlives the session, ZooKeeper expires it, the master starts recovering the server's regions, and the RegionServer aborts when it notices. Raising the timeout hides the problem and slows down recovery from real crashes. The durable fix is to keep the worst-case pause well below the timeout, and that is the whole goal of GC tuning for HBase: bound the tail, not the average. The RegionServer architecture article covers the recovery path in detail.

What the heap holds, and for how long

Collectors are designed around the assumption that most objects die young. A RegionServer has three populations that behave very differently.

RegionServer memory: what lives where, and how long it livesJava heap (-Xms = -Xmx, for example 31 GB)MemStore (0.4 of heap)MSLAB 2 MB chunks, lives until flushL1 block cacheindex and bloom blocksRPC and request objectsdie within millisecondsScanners, compactionshort to medium livedG1 young and old regionsshort-lived garbage dies young; MemStore and L1 cache promote to oldOff-heap (direct memory)BucketCache data blockssized by -XX:MaxDirectMemorySize(HBASE_OFFHEAPSIZE)Off-heap RPC bufferspooled, reusedOS page cache, DataNodeleave room for themZooKeeper session: a pause longer than the timeoutends the session; the master reassigns regions; the RegionServer aborts
The heap holds MemStore, the L1 cache and short-lived request objects; large data caches belong off-heap in direct memory, which the collector does not scan or move.
PopulationLifetimeEffect on the collector
RPC requests, result objects, iteratorsMillisecondsCheap: die in the young generation
MemStore cellsSeconds to minutes, until flushPromoted to old, then all die together at flush
On-heap block cache blocksUntil evicted, often longOld generation, churn under eviction

MemStore is the awkward one: it fills the old generation gradually and then releases large amounts at once when a region flushes, with blocks of cache being evicted at the same time. On collectors that do not compact the old generation, this leaves the heap fragmented: plenty of free space, none of it contiguous. Eventually a promotion cannot find room and the JVM falls back to a stop-the-world full collection that compacts the whole heap. On a 30 GB heap that can take tens of seconds, which is exactly the session-killing pause. The flush mechanics are covered in MemStore flushing.

Advertisement

MSLAB: allocating MemStore in chunks

The MemStore-Local Allocation Buffer fixes fragmentation at the source. Instead of allocating each cell as its own small object, HBase copies cells into 2 MB chunks (hbase.hregion.memstore.mslab.chunksize = 2097152) that belong to one region's MemStore. When the region flushes, whole chunks become garbage at once, so the heap sees a few large uniform objects coming and going rather than millions of small ones with mixed lifetimes. Cells larger than 256 KB (mslab.max.allocation = 262144) bypass the chunks and are allocated directly. MSLAB is enabled by default, and you should leave it on.

The chunk pool (hbase.hregion.memstore.chunkpool.maxsize) goes one step further by recycling chunks after a flush instead of freeing them, which reduces allocation work. The cost of MSLAB is that each region with a MemStore holds at least one chunk, so thousands of regions per server waste memory in partly filled chunks; that is another reason to keep region counts moderate.

Which collector actually runs

The shipped hbase-env.sh sets no collector: the GC lines in it are commented-out logging examples. So an unconfigured RegionServer uses the JVM default, which is the Parallel collector on JDK 8 and G1 on JDK 9 and later. Parallel GC performs every old-generation collection as a stop-the-world pause proportional to live data, which is unacceptable for a large RegionServer. Older HBase guidance recommended CMS, but CMS was deprecated in JDK 9 and removed in JDK 14, so it is not an option on the JDK 17 that current releases support.

In practice that leaves G1 as the default choice. It divides the heap into equal regions, collects young regions in short parallel pauses, marks the old generation concurrently, and evacuates the old regions with the most garbage in mixed collections, compacting as it goes, so fragmentation does not accumulate the way it does with CMS. ZGC and Shenandoah do almost all their work concurrently and keep pauses in the low milliseconds regardless of heap size. They are worth testing only on a JDK and HBase version you have validated together; generational ZGC in particular needs JDK 21, which is newer than the JDK versions the HBase book lists as supported. See the G1 collector and ZGC for the collectors themselves.

G1 settings that matter for a RegionServer

# hbase-env.sh (JDK 11 or 17, G1)
export HBASE_HEAPSIZE=31G
export HBASE_OFFHEAPSIZE=70G     # becomes -XX:MaxDirectMemorySize
export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS \
  -Xms31g -XX:+UseG1GC \
  -XX:MaxGCPauseMillis=100 \
  -XX:G1HeapRegionSize=32m \
  -XX:InitiatingHeapOccupancyPercent=45 \
  -XX:+ParallelRefProcEnabled \
  -XX:+AlwaysPreTouch \
  -Xlog:gc*,safepoint:file=/var/log/hbase/gc-rs.log:time,uptime,level,tags:filecount=10,filesize=100m"
  • Fixed heap. HBASE_HEAPSIZE sets the maximum heap; add a matching -Xms so the heap never resizes; AlwaysPreTouch touches every page at start-up, so the first flush does not stall on page faults. Start-up takes longer as a result.
  • MaxGCPauseMillis. A target, not a guarantee. G1 shrinks the young generation to meet it. Setting it very low makes young collections more frequent and promotes more objects prematurely; 100 to 200 ms is a reasonable range to begin measuring from. The default is 200.
  • Region size and humongous objects. Any object of at least half a G1 region is humongous: it is allocated directly into contiguous old regions, which is expensive and can trigger collections. A 2 MB MSLAB chunk with 4 MB regions is humongous; with 8 MB or larger regions it is not. Large cells, big multi-gets and large scan results create humongous arrays too. Look for humongous allocations in the log and raise the region size if they are frequent.
  • Marking threshold. InitiatingHeapOccupancyPercent sets when concurrent marking starts; modern G1 adapts it at run time and uses this value as the starting point. If marking finishes too late, G1 runs out of free regions and fails evacuation, which leads to long pauses; start marking earlier rather than later.
  • ParallelRefProcEnabled. Processes soft, weak and phantom references in parallel, which helps when reference processing shows up as a long phase in pauses.

Change one setting at a time and measure pause percentiles under a realistic write and read load, including flushes and compactions. A setting that improves a synthetic read test can make pauses worse under a flush storm.

Heap size and moving the cache off-heap

Keep the heap below about 32 GB. Up to that size the JVM can use compressed object pointers, so references take four bytes instead of eight; a 40 GB heap can hold less data than a 31 GB one. More importantly, every gigabyte of heap is something the collector must mark and eventually evacuate, so a larger heap raises the worst case.

The way to use a large machine is to move bulk data out of the heap. BucketCache with hbase.bucketcache.ioengine=offheap stores data blocks in direct memory that G1 never scans or moves, while the on-heap L1 cache keeps the small, hot index and bloom blocks. Set hbase.bucketcache.size in megabytes, and set HBASE_OFFHEAPSIZE (the direct-memory limit) larger than the BucketCache so the off-heap RPC buffers also fit. HBase 2.x can also place MemStore chunks off-heap with hbase.regionserver.offheap.global.memstore.size, which defaults to 0 (disabled); it helps write-heavy servers but makes the direct-memory budget more important. The eviction logic of the caches is covered in the block cache.

HBase also checks that MemStore plus on-heap block cache do not exceed 80% of the heap, and refuses to start otherwise. With the defaults of 0.4 each you are exactly at that limit, which leaves only 20% for everything else. Once data blocks live in BucketCache, lower hfile.block.cache.size so the L1 cache only needs to hold index and bloom blocks, and give the headroom back.

Reading a GC log

Enable unified logging as in the example and keep the logs. The lines that matter in a G1 log start with Pause Young (Normal), Pause Young (Concurrent Start), Pause Young (Mixed), Pause Remark, Pause Cleanup and, the one you never want, Pause Full. Each reports heap before and after and the pause length. Collect the pauses, not the averages.

import re, sys
pat = re.compile(r"Pause (Young|Remark|Cleanup|Full)[^\n]*?([\d.]+)ms")
pauses = {}
for line in open(sys.argv[1], encoding="utf-8", errors="replace"):
    m = pat.search(line)
    if m:
        pauses.setdefault(m.group(1), []).append(float(m.group(2)))
for kind, xs in sorted(pauses.items()):
    xs.sort()
    p99 = xs[min(len(xs) - 1, int(len(xs) * 0.99))]
    print(f"{kind:8} n={len(xs):6} p99={p99:8.1f}ms max={xs[-1]:8.1f}ms")

Four patterns explain most bad logs. A Pause Full after an evacuation failure (reported as to-space exhausted in older JDKs) means G1 ran out of free regions: start marking earlier, reduce live data on heap, or relax an overly aggressive pause target. Frequent humongous allocations mean the region size is too small for your objects. Young pauses that grow steadily usually mean too much is surviving, often a cache or MemStore too large for the heap. Pauses reported by HBase's own JVM pause monitor with no matching GC entry point to a cause outside the JVM: swapping, an overloaded host or a stalled disk. Disable swap on HBase hosts or set swappiness to a minimum.

Worked example: budgeting a 128 GB server

ConsumerSizeReasoning
Java heap31 GBBelow the compressed-pointer limit
MemStore (0.4 of heap)about 12.4 GBWrite buffer; flush size times busy regions must fit
L1 block cache (0.1 of heap)about 3.1 GBIndex and bloom blocks only
BucketCache off-heap64 GBhbase.bucketcache.size = 65536 (MB)
Direct memory limit70 GBBucketCache plus about 6 GB for RPC buffers
DataNode, OS, page cacheabout 27 GBThe remainder; never give it all away

MemStore plus L1 is 0.5 of the heap, comfortably under the 80% check, and leaves about 15.5 GB of heap for request objects, scanners and compaction. Total committed memory is 31 + 70 = 101 GB, leaving the operating system, the co-located DataNode and the page cache 27 GB. Run a write load that forces flushes and a read load that churns the BucketCache, record pause percentiles, and adjust from measurements, not from this table.

Failure modes

SymptomCauseFix
RegionServer aborts after a session expiryPause longer than the effective ZooKeeper timeoutFind the pause in the GC log; reduce heap live data; check the ensemble maxSessionTimeout
Pause Full after evacuation failureMarking started too late or the heap is too fullLower IHOP, reduce on-heap cache, leave more free heap
Humongous allocationsG1 region too small for chunks, cells or resultsRaise G1HeapRegionSize; cap result sizes
OutOfMemoryError: Direct buffer memoryDirect memory limit below BucketCache plus buffersRaise HBASE_OFFHEAPSIZE or shrink BucketCache
Long pauses without GC entriesSwap, host overload or I/O stallsDisable swap, check steal time and disk latency
Start-up failure on the 80% checkMemStore plus block cache fractions too largeLower hfile.block.cache.size once BucketCache is used

What to do next

  1. Check which JDK and collector each RegionServer actually runs; do not assume hbase-env.sh sets one.
  2. Enable unified GC logging with rotation on every RegionServer and compute p99 and maximum pauses by type.
  3. Look up the ZooKeeper ensemble's maxSessionTimeout and write down the effective session timeout.
  4. Fix the heap below 32 GB, switch to G1 with a measured pause target, and check for humongous allocations.
  5. Move data blocks to an off-heap BucketCache, lower the on-heap cache fraction, and size direct memory with headroom.
  6. Leave MSLAB on and keep the region count per server moderate.
  7. Load-test with flushes and compactions running, then change one setting at a time.
Key takeaway: HBase GC tuning is about the worst pause, because a pause longer than the effective ZooKeeper session timeout makes a RegionServer abort. Know which collector actually runs, keep MSLAB on, use G1 with a fixed heap under 32 GB, avoid humongous allocations, and put the bulk data cache off-heap in BucketCache with a correctly sized direct-memory limit. Then read the GC logs by pause type and tune one setting at a time under realistic load.