A RegionServer that stops for a long garbage-collection pause does not just answer slowly. If the pause outlasts its ZooKeeper session, the master declares it dead, reassigns its regions and replays its write-ahead log, and the server discovers on waking that it no longer owns anything. So the collector choice for HBase is not a micro-optimisation: it decides the worst case of the whole cluster.

The HBase JVM GC tuning article covers heap layout, MSLAB, a full G1 flag set and how to read G1 logs. This article is about the decision before that one: which collector to run at all. It explains what a RegionServer demands from a collector, how G1, ZGC and Shenandoah meet those demands and how each fails, what the JDK support matrix allows, and how to run an A/B test on real traffic so the choice rests on your own measurements rather than on benchmarks from someone else's heap.

Advertisement

What a RegionServer asks of a collector

A RegionServer's heap holds three populations with very different lifetimes. RPC request and response objects, cell copies and iterator state live for milliseconds and die young. MemStore data lives from the write until the next flush, seconds to minutes, which is too long for the young generation and too short to be truly old; MSLAB allocates it in fixed-size chunks so it does not fragment the old generation. An on-heap block cache holds blocks for as long as they stay hot and is evicted in bulk, which is steady churn in the old generation.

A good collector for this mix does three things. It keeps the worst-case pause far below the ZooKeeper session timeout and, ideally, below the client's RPC timeout. It keeps up with the allocation rate at peak write load without falling back to a full stop-the-world collection. And it leaves enough CPU for the handler threads, because a collector that runs concurrently still uses cores the RegionServer would otherwise spend serving requests. The RegionServer article describes where those allocations come from.

Three collectors, three designs

G1 is the default on JDK 9 and later. It divides the heap into equal regions, collects the young generation in stop-the-world evacuation pauses, marks the old generation concurrently, and then evacuates the old regions with the most garbage in mixed collections. Its pause time scales with the amount of live data it copies per pause, so a large young generation or a large on-heap cache pushes pauses up. You steer it with a pause-time target that it tries to meet by resizing the young generation.

ZGC does marking and relocation concurrently with the application, using coloured pointers and load barriers so that threads fix up references to moved objects as they read them. Its pauses are short and do not grow with heap size. The costs are CPU spent in barriers and concurrent threads, and memory: ZGC does not use compressed object pointers, so every reference takes 8 bytes instead of 4 and the same data needs more heap. Generational ZGC arrived in JDK 21 (JEP 439), became the default ZGC mode in JDK 23 (JEP 474), and JDK 24 removed the single-generation mode entirely (JEP 490).

Shenandoah also evacuates concurrently, using load-reference barriers, and supports compressed pointers. It uses pacing: when allocation outruns the concurrent cycle, it slows allocating threads down a little rather than stopping everything. Its generational mode was added as experimental in JDK 24 (JEP 404) and became a product feature in JDK 25 (JEP 521). The ZGC architecture article and the Shenandoah article explain the barrier mechanics.

G1ZGCShenandoah
Pause sourceYoung and mixed evacuationShort root-scan pauses onlyShort init and final pauses
Pause vs heap sizeGrows with live data copiedRoughly flatRoughly flat
Extra costRemembered sets, write barriersLoad barriers, concurrent threads, no compressed oopsLoad-reference barriers, concurrent threads
How it fails under pressureTo-space exhausted, then full GCAllocation stall: threads wait for the cyclePacing, then degenerated GC, then full GC
GenerationalAlwaysJDK 21+; only mode from JDK 24Experimental JDK 24, product JDK 25
Advertisement

What the JDK support matrix allows

The HBase reference guide's Java support matrix is the constraint that most narrows the choice. It states that JDK 17 support was introduced with HBase 2.5.x, and at the time of writing its matrix does not list JDK 21 or later. The book also recommends long-term-support JDKs only. On a supported JDK 17, then, your realistic options are G1, single-generation ZGC, and Shenandoah if your JDK build includes it. Not every vendor ships Shenandoah, so check with java -XX:+UseShenandoahGC -version on the exact JDK your RegionServers run.

Generational ZGC, which removes single-generation ZGC's biggest weakness, needs JDK 21 or later, and generational Shenandoah as a product feature needs JDK 25. Running HBase on those JDKs means operating outside the documented matrix: possible, and some operators do it, but you own the compatibility testing, including Hadoop client libraries, coprocessors and any agents attached to the JVM.

Which JDK can this cluster run?HBase book matrix: newest listed is 17JDK 17 (supported)G1, single-gen ZGC, Shenandoah*JDK 21 / 25 (unlisted)generational ZGC; gen. Shenandoah on 25Shrink the heap firstBucketCache off-heap, MSLAB onG1 (default)pauses scale with young genZGCsub-ms pauses, more CPU, no compressed oopsShenandoahconcurrent evacuation, pacing* only if your JDK build includes it: check with java -XX:+UseShenandoahGC -version
Collector choice for a RegionServer. The JDK you can run decides the options; shrinking the heap by moving the block cache off-heap comes before any collector change; then an A/B test decides.

Shrink the heap before you switch collectors

The largest pause reduction usually comes from making the heap smaller, not from changing the collector. An on-heap block cache of tens of gigabytes is long-lived data that G1 must mark and eventually move, and that any concurrent collector must trace every cycle. Moving the cache off-heap with BucketCache in off-heap mode takes that data out of the collector's reach entirely; the BucketCache article explains the configuration and its costs. With MSLAB on and the cache off-heap, many RegionServers run comfortably with G1 on a heap of 16 to 32 GB.

This step also changes the comparison. Single-generation ZGC's weakness is that every cycle marks the whole live heap, so its CPU cost and its risk of allocation stalls grow with the live set. A small heap with a modest live set is exactly where single-generation ZGC on JDK 17 does well. A 31 GB heap full of on-heap cache is where it struggles. Measure after the heap change, not before.

Collector-specific configuration

Keep everything except the collector identical between arms of the test: heap size policy, logging, MSLAB and cache settings. The one deliberate exception is ZGC's heap size. Because it has no compressed oops, the same live data takes more room, so give the ZGC arm somewhat more heap and check that the machine's memory budget, including off-heap cache and the operating system page cache, still fits.

# hbase-env.sh -- one block per RegionServer group under test. Keep heap, logging
# and everything else identical; change only the collector lines.
GC_LOG="-Xlog:gc*,safepoint:file=/var/log/hbase/gc-rs.log:time,uptime,level,tags:filecount=10,filesize=100m"

# Control: G1 (see the G1 tuning article for the full flag set)
export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS -Xms24g -Xmx24g -XX:+UseG1GC -XX:MaxGCPauseMillis=100 $GC_LOG"

# Candidate A: ZGC. On JDK 17 this is single-generation ZGC; on JDK 21 and 22 add
# -XX:+ZGenerational; from JDK 23 generational is the default mode, and JDK 24
# removed the non-generational mode.
# export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS -Xms28g -Xmx28g -XX:+UseZGC -XX:+AlwaysPreTouch $GC_LOG"

# Candidate B: Shenandoah (only if 'java -XX:+UseShenandoahGC -version' succeeds).
# JDK 24 needs -XX:+UnlockExperimentalVMOptions -XX:ShenandoahGCMode=generational
# for the generational mode; JDK 25 made it a product option.
# export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS -Xms24g -Xmx24g -XX:+UseShenandoahGC $GC_LOG"

Two more settings matter for concurrent collectors. They need spare cores: if the RegionServer already runs its CPUs near saturation at peak, concurrent GC threads will steal from handlers and latency will rise even though pauses fall. And -XX:+AlwaysPreTouch is worth using on all arms, so that page faults on a freshly grown heap do not appear as GC or latency noise during the test.

An A/B test that produces an answer

Pick two RegionServers on identical hardware with comparable region load, or two small groups using RegionServer groups so that the regions are fixed. Give one the control collector and the other the candidate. Run both through at least one full daily cycle of production traffic, including compactions and the heaviest write period, or drive them with a YCSB workload shaped like your production read, write and scan mix.

Collect three kinds of evidence. From the GC log: the pause distribution up to p99.9 and the maximum, plus any allocation stalls, degenerated cycles, full collections or to-space exhaustion. From HBase: the RegionServer's latency histograms for gets, puts and scans, exposed over JMX and the metrics system, and the count of slow-operation warnings. From the host: CPU utilisation and run-queue length at peak. The parser below extracts the first set from unified JVM logs for any of the three collectors.

import re, sys, statistics

# Unified-logging pause lines end in a duration such as "12.345ms". G1 logs them
# under [gc], ZGC under [gc,phases] (generational ZGC adds a "Y:" or "O:" prefix):
#   [..][info][gc          ] GC(812) Pause Young (Normal) (G1 Evacuation Pause) 9012M->2210M(24576M) 18.402ms
#   [..][info][gc,phases   ] GC(0) Y: Pause Mark Start (Major) 0.034ms
PAUSE = re.compile(r"\[gc(?:,phases)?\s*\] GC\(\d+\) ((?:[yYO]: )?Pause .*?)\s+"
                   r"(?:\d+M->\d+M\(\d+M\)\s+)?(\d+\.\d+)ms$")
# "Allocation Stall (" is a real stall; ZGC's per-cycle "Allocation Stalls:" table is not.
TROUBLE = ("Allocation Stall (", "Degenerated GC", "Pause Full", "To-space exhausted")

def summarise(path):
    pauses, trouble = {}, {t: 0 for t in TROUBLE}
    for line in open(path, errors="replace"):
        if "[gc,start" in line:              # G1 announces each pause twice
            continue
        m = PAUSE.search(line.rstrip())
        if m:
            pauses.setdefault(m.group(1).strip(), []).append(float(m.group(2)))
        for t in TROUBLE:
            trouble[t] += t in line
    every = sorted(v for vals in pauses.values() for v in vals)
    q = lambda x: every[min(len(every) - 1, int(x * len(every)))] if every else 0.0
    print(f"{path}: {len(every)} pauses, p50 {q(.5):.1f} ms, p99 {q(.99):.1f} ms, "
          f"p99.9 {q(.999):.1f} ms, max {max(every, default=0):.1f} ms")
    for kind, vals in sorted(pauses.items()):
        print(f"  {kind:40s} n={len(vals):6d} mean={statistics.mean(vals):7.2f} ms")
    print("  trouble:", trouble)

for f in sys.argv[1:]:
    summarise(f)

Decide on client-visible tail latency and on the absence of trouble events, not on the mean pause. A collector that cuts p99.9 pause from 300 ms to 2 ms but raises p99 get latency by 10 percent because handlers lost CPU may still be the right choice if your SLA is written at p99.9, and the wrong one if it is written at p99. Write the decision criterion down before the test starts.

Worked example: from 31 GB G1 to a smaller heap and ZGC

A read-heavy cluster on JDK 17 runs RegionServers with a 31 GB G1 heap, 12 GB of it on-heap block cache. The figures that follow are illustrative. The GC log parser shows p99 pauses of 140 ms and a p99.9 of 420 ms, with occasional mixed-collection pauses above 600 ms during compaction storms, and client p99.9 get latency tracks them.

Step one moves the block cache to off-heap BucketCache and cuts the heap to 20 GB. G1's p99.9 pause falls to 110 ms, because the old generation now holds MemStore chunks and little else. Step two is the A/B: one RegionServer group keeps G1 at 20 GB, the other runs ZGC at 24 GB to cover the loss of compressed oops. Over a week, ZGC's maximum pause stays under 2 ms and client p99.9 get latency drops by about a third. CPU at peak rises from 55 to 63 percent, and the log shows no allocation stalls. The team adopts ZGC for the read-heavy group, keeps G1 on a write-heavy group whose CPU is already above 80 percent at peak, and schedules a re-test when the cluster moves to a JDK that HBase lists as supported with generational ZGC.

Failure modes

  • Switching collectors to fix an oversized heap. Move the cache off-heap first; the collector change may then be unnecessary.
  • Starving handlers. Concurrent collectors use cores. On a CPU-bound RegionServer, pauses fall and request latency rises.
  • Ignoring allocation stalls. A ZGC stall pauses the allocating thread, not the whole JVM, so it never appears as a GC pause. Count stalls separately.
  • Same heap size for ZGC. Without compressed oops, a heap sized for G1 is effectively smaller under ZGC.
  • Testing on quiet traffic. Collectors differ most under compaction and write peaks. Include them.
  • Running an unlisted JDK without owning the testing. Coprocessors, agents and Hadoop libraries may break before the collector does.

What to do next

  1. Run the log parser over a week of your current GC logs and record pause p99, p99.9, maximum and trouble events.
  2. Check your ZooKeeper session and client RPC timeouts against the maximum pause you found.
  3. If the block cache is on-heap, move it off-heap and shrink the heap, then re-measure.
  4. Confirm which collectors your exact JDK build supports, including Shenandoah, before planning a test.
  5. Run a one-week A/B with written decision criteria on tail latency, trouble events and peak CPU.
  6. Adopt per RegionServer group where workloads differ, and repeat the test after every JDK upgrade.
Key takeaway: Choosing an HBase collector starts with the JDK you can support: on JDK 17 that means G1, single-generation ZGC, or Shenandoah if your build ships it; generational ZGC needs JDK 21 or later, outside HBase's documented matrix. Shrink the heap by moving the block cache off-heap before anything else. Then let a controlled A/B on real traffic decide, judging tail latency, allocation stalls and full collections, and the CPU a concurrent collector takes from your handlers.