In most Java services a long garbage-collection pause is a latency problem. In HBase it is an availability problem. A RegionServer that stops for longer than its ZooKeeper session timeout is declared dead, its regions are reassigned to other servers, and when it wakes up it discovers it has been fenced off and aborts. A single bad pause can therefore take a server out of the cluster and trigger a wave of region moves and WAL splitting that hurts every other server too.
This article explains GC tuning for HBase 2.x RegionServers from the allocation pattern upwards: what lives in the heap and for how long, how MSLAB prevents fragmentation, which G1 settings matter and why, when to move data off-heap, how to read a GC log, and how to budget a real server's memory. It ends with a failure table and a checklist. Defaults quoted come from hbase-default.xml on the 2.6 branch; verify them against the version you run.
Why a pause looks like death
A RegionServer holds a session with ZooKeeper and keeps it alive with heartbeats. HBase sets zookeeper.session.timeout to 90,000 ms by default, but that is only a request: the ZooKeeper server caps sessions at its own maxSessionTimeout, which defaults to 20 times its tick time, or 40 seconds with the usual 2-second tick. Unless the ensemble has been configured to allow more, your real timeout is 40 seconds, not 90. This mismatch surprises many operators.
When the JVM stops every thread for a full collection, heartbeats stop too. If the pause outlives the session, ZooKeeper expires it, the master starts recovering the server's regions, and the RegionServer aborts when it notices. Raising the timeout hides the problem and slows down recovery from real crashes. The durable fix is to keep the worst-case pause well below the timeout, and that is the whole goal of GC tuning for HBase: bound the tail, not the average. The RegionServer architecture article covers the recovery path in detail.
What the heap holds, and for how long
Collectors are designed around the assumption that most objects die young. A RegionServer has three populations that behave very differently.
| Population | Lifetime | Effect on the collector |
|---|---|---|
| RPC requests, result objects, iterators | Milliseconds | Cheap: die in the young generation |
| MemStore cells | Seconds to minutes, until flush | Promoted to old, then all die together at flush |
| On-heap block cache blocks | Until evicted, often long | Old generation, churn under eviction |
MemStore is the awkward one: it fills the old generation gradually and then releases large amounts at once when a region flushes, with blocks of cache being evicted at the same time. On collectors that do not compact the old generation, this leaves the heap fragmented: plenty of free space, none of it contiguous. Eventually a promotion cannot find room and the JVM falls back to a stop-the-world full collection that compacts the whole heap. On a 30 GB heap that can take tens of seconds, which is exactly the session-killing pause. The flush mechanics are covered in MemStore flushing.
MSLAB: allocating MemStore in chunks
The MemStore-Local Allocation Buffer fixes fragmentation at the source. Instead of allocating each cell as its own small object, HBase copies cells into 2 MB chunks (hbase.hregion.memstore.mslab.chunksize = 2097152) that belong to one region's MemStore. When the region flushes, whole chunks become garbage at once, so the heap sees a few large uniform objects coming and going rather than millions of small ones with mixed lifetimes. Cells larger than 256 KB (mslab.max.allocation = 262144) bypass the chunks and are allocated directly. MSLAB is enabled by default, and you should leave it on.
The chunk pool (hbase.hregion.memstore.chunkpool.maxsize) goes one step further by recycling chunks after a flush instead of freeing them, which reduces allocation work. The cost of MSLAB is that each region with a MemStore holds at least one chunk, so thousands of regions per server waste memory in partly filled chunks; that is another reason to keep region counts moderate.
Which collector actually runs
The shipped hbase-env.sh sets no collector: the GC lines in it are commented-out logging examples. So an unconfigured RegionServer uses the JVM default, which is the Parallel collector on JDK 8 and G1 on JDK 9 and later. Parallel GC performs every old-generation collection as a stop-the-world pause proportional to live data, which is unacceptable for a large RegionServer. Older HBase guidance recommended CMS, but CMS was deprecated in JDK 9 and removed in JDK 14, so it is not an option on the JDK 17 that current releases support.
In practice that leaves G1 as the default choice. It divides the heap into equal regions, collects young regions in short parallel pauses, marks the old generation concurrently, and evacuates the old regions with the most garbage in mixed collections, compacting as it goes, so fragmentation does not accumulate the way it does with CMS. ZGC and Shenandoah do almost all their work concurrently and keep pauses in the low milliseconds regardless of heap size. They are worth testing only on a JDK and HBase version you have validated together; generational ZGC in particular needs JDK 21, which is newer than the JDK versions the HBase book lists as supported. See the G1 collector and ZGC for the collectors themselves.
G1 settings that matter for a RegionServer
# hbase-env.sh (JDK 11 or 17, G1)
export HBASE_HEAPSIZE=31G
export HBASE_OFFHEAPSIZE=70G # becomes -XX:MaxDirectMemorySize
export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS \
-Xms31g -XX:+UseG1GC \
-XX:MaxGCPauseMillis=100 \
-XX:G1HeapRegionSize=32m \
-XX:InitiatingHeapOccupancyPercent=45 \
-XX:+ParallelRefProcEnabled \
-XX:+AlwaysPreTouch \
-Xlog:gc*,safepoint:file=/var/log/hbase/gc-rs.log:time,uptime,level,tags:filecount=10,filesize=100m"- Fixed heap. HBASE_HEAPSIZE sets the maximum heap; add a matching -Xms so the heap never resizes; AlwaysPreTouch touches every page at start-up, so the first flush does not stall on page faults. Start-up takes longer as a result.
- MaxGCPauseMillis. A target, not a guarantee. G1 shrinks the young generation to meet it. Setting it very low makes young collections more frequent and promotes more objects prematurely; 100 to 200 ms is a reasonable range to begin measuring from. The default is 200.
- Region size and humongous objects. Any object of at least half a G1 region is humongous: it is allocated directly into contiguous old regions, which is expensive and can trigger collections. A 2 MB MSLAB chunk with 4 MB regions is humongous; with 8 MB or larger regions it is not. Large cells, big multi-gets and large scan results create humongous arrays too. Look for humongous allocations in the log and raise the region size if they are frequent.
- Marking threshold. InitiatingHeapOccupancyPercent sets when concurrent marking starts; modern G1 adapts it at run time and uses this value as the starting point. If marking finishes too late, G1 runs out of free regions and fails evacuation, which leads to long pauses; start marking earlier rather than later.
- ParallelRefProcEnabled. Processes soft, weak and phantom references in parallel, which helps when reference processing shows up as a long phase in pauses.
Change one setting at a time and measure pause percentiles under a realistic write and read load, including flushes and compactions. A setting that improves a synthetic read test can make pauses worse under a flush storm.
Heap size and moving the cache off-heap
Keep the heap below about 32 GB. Up to that size the JVM can use compressed object pointers, so references take four bytes instead of eight; a 40 GB heap can hold less data than a 31 GB one. More importantly, every gigabyte of heap is something the collector must mark and eventually evacuate, so a larger heap raises the worst case.
The way to use a large machine is to move bulk data out of the heap. BucketCache with hbase.bucketcache.ioengine=offheap stores data blocks in direct memory that G1 never scans or moves, while the on-heap L1 cache keeps the small, hot index and bloom blocks. Set hbase.bucketcache.size in megabytes, and set HBASE_OFFHEAPSIZE (the direct-memory limit) larger than the BucketCache so the off-heap RPC buffers also fit. HBase 2.x can also place MemStore chunks off-heap with hbase.regionserver.offheap.global.memstore.size, which defaults to 0 (disabled); it helps write-heavy servers but makes the direct-memory budget more important. The eviction logic of the caches is covered in the block cache.
HBase also checks that MemStore plus on-heap block cache do not exceed 80% of the heap, and refuses to start otherwise. With the defaults of 0.4 each you are exactly at that limit, which leaves only 20% for everything else. Once data blocks live in BucketCache, lower hfile.block.cache.size so the L1 cache only needs to hold index and bloom blocks, and give the headroom back.
Reading a GC log
Enable unified logging as in the example and keep the logs. The lines that matter in a G1 log start with Pause Young (Normal), Pause Young (Concurrent Start), Pause Young (Mixed), Pause Remark, Pause Cleanup and, the one you never want, Pause Full. Each reports heap before and after and the pause length. Collect the pauses, not the averages.
import re, sys
pat = re.compile(r"Pause (Young|Remark|Cleanup|Full)[^\n]*?([\d.]+)ms")
pauses = {}
for line in open(sys.argv[1], encoding="utf-8", errors="replace"):
m = pat.search(line)
if m:
pauses.setdefault(m.group(1), []).append(float(m.group(2)))
for kind, xs in sorted(pauses.items()):
xs.sort()
p99 = xs[min(len(xs) - 1, int(len(xs) * 0.99))]
print(f"{kind:8} n={len(xs):6} p99={p99:8.1f}ms max={xs[-1]:8.1f}ms")Four patterns explain most bad logs. A Pause Full after an evacuation failure (reported as to-space exhausted in older JDKs) means G1 ran out of free regions: start marking earlier, reduce live data on heap, or relax an overly aggressive pause target. Frequent humongous allocations mean the region size is too small for your objects. Young pauses that grow steadily usually mean too much is surviving, often a cache or MemStore too large for the heap. Pauses reported by HBase's own JVM pause monitor with no matching GC entry point to a cause outside the JVM: swapping, an overloaded host or a stalled disk. Disable swap on HBase hosts or set swappiness to a minimum.
Worked example: budgeting a 128 GB server
| Consumer | Size | Reasoning |
|---|---|---|
| Java heap | 31 GB | Below the compressed-pointer limit |
| MemStore (0.4 of heap) | about 12.4 GB | Write buffer; flush size times busy regions must fit |
| L1 block cache (0.1 of heap) | about 3.1 GB | Index and bloom blocks only |
| BucketCache off-heap | 64 GB | hbase.bucketcache.size = 65536 (MB) |
| Direct memory limit | 70 GB | BucketCache plus about 6 GB for RPC buffers |
| DataNode, OS, page cache | about 27 GB | The remainder; never give it all away |
MemStore plus L1 is 0.5 of the heap, comfortably under the 80% check, and leaves about 15.5 GB of heap for request objects, scanners and compaction. Total committed memory is 31 + 70 = 101 GB, leaving the operating system, the co-located DataNode and the page cache 27 GB. Run a write load that forces flushes and a read load that churns the BucketCache, record pause percentiles, and adjust from measurements, not from this table.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| RegionServer aborts after a session expiry | Pause longer than the effective ZooKeeper timeout | Find the pause in the GC log; reduce heap live data; check the ensemble maxSessionTimeout |
| Pause Full after evacuation failure | Marking started too late or the heap is too full | Lower IHOP, reduce on-heap cache, leave more free heap |
| Humongous allocations | G1 region too small for chunks, cells or results | Raise G1HeapRegionSize; cap result sizes |
| OutOfMemoryError: Direct buffer memory | Direct memory limit below BucketCache plus buffers | Raise HBASE_OFFHEAPSIZE or shrink BucketCache |
| Long pauses without GC entries | Swap, host overload or I/O stalls | Disable swap, check steal time and disk latency |
| Start-up failure on the 80% check | MemStore plus block cache fractions too large | Lower hfile.block.cache.size once BucketCache is used |
What to do next
- Check which JDK and collector each RegionServer actually runs; do not assume hbase-env.sh sets one.
- Enable unified GC logging with rotation on every RegionServer and compute p99 and maximum pauses by type.
- Look up the ZooKeeper ensemble's maxSessionTimeout and write down the effective session timeout.
- Fix the heap below 32 GB, switch to G1 with a measured pause target, and check for humongous allocations.
- Move data blocks to an off-heap BucketCache, lower the on-heap cache fraction, and size direct memory with headroom.
- Leave MSLAB on and keep the region count per server moderate.
- Load-test with flushes and compactions running, then change one setting at a time.