Moving a RegionServer's block cache and buffers off the Java heap shortens garbage-collection pauses, but it also moves the memory where the usual Java tools stop looking. A heap dump no longer shows the cache. GC logs no longer show the buffers. When the process grows past its budget, or the kernel kills it, the question is no longer which objects are on the heap but which layer outside the heap grew.
This article is about operating off-heap memory once it is configured: how to account for every gigabyte of the process's resident memory, how to find reference-count leaks in HBase's pooled buffers and native leaks outside them, and how to roll off-heap settings out across a cluster with a canary and a rollback. The budget itself, including BucketCache, ByteBuffAllocator defaults and a full sizing worksheet, is covered in the off-heap memory budget article; this one assumes you have a budget and need to prove the process stays inside it.
The layers of RegionServer memory
The kernel sees one number per process, its resident set size (RSS), and a container limit or the OOM killer acts on that number alone. Inside a RegionServer, RSS is made of layers, each with its own owner and its own meter.
- Committed Java heap. Up to
-Xmx, set byHBASE_HEAPSIZE. With-Xmsequal to-Xmxand pre-touch it is all resident from start-up; otherwise it grows. - Direct buffers. Capped by
-XX:MaxDirectMemorySize, set fromHBASE_OFFHEAPSIZE. This holds the off-heap BucketCache, the ByteBuffAllocator pool, off-heap MemStore chunks if enabled, and network and HDFS client buffers. - JVM native memory. Metaspace, the code cache, thread stacks, and garbage collector bookkeeping, which for G1 can be a few percent of the heap. None of this is capped by either flag above.
- Native allocations outside the JVM. Native compression libraries and anything else calling
malloc, plus fragmentation in the C allocator itself.
Two of these layers have hard caps that the JVM enforces by throwing an error: the heap and direct memory. The other two have no cap at all inside the process. They are limited only by the host or the container, and the kernel enforces that limit by killing the process without a Java stack trace. That asymmetry shapes everything below. A budget that adds up the two capped layers and forgets the uncapped ones looks correct on paper and fails in production, usually weeks after the change, when a compaction storm or a traffic peak pushes native use up.
Note also the difference between reserved and committed memory. The JVM reserves address space generously, for example for the whole heap and for metaspace, but only committed pages that have been touched count towards RSS. Always compare RSS with committed figures, never with reserved ones.
The direct-memory default trap
One default deserves its own section because it breaks budgets silently. If -XX:MaxDirectMemorySize is not set, the JVM's direct-memory limit defaults to the maximum heap size. A RegionServer started without HBASE_OFFHEAPSIZE therefore has a direct limit equal to -Xmx, and an off-heap BucketCache configured larger than that fails at start-up or, worse, a cache that fits leaves no room for the allocator pool.
The opposite surprise is just as common. With a 31 GB heap and no explicit limit, the process may legitimately use 31 GB of heap plus up to 31 GB of direct memory, which a container sized only for the heap will kill. Always set the direct limit explicitly, and size the container or host budget for heap plus direct limit plus native overhead.
Reconciling RSS with Native Memory Tracking
Reconciliation means comparing RSS with the sum of what the meters explain, layer by layer. Turn on Native Memory Tracking (NMT) in summary mode on the RegionServer you are investigating. It adds a small overhead, so keep it to the nodes under study.
# hbase-env.sh on the canary RegionServer only
export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS -XX:NativeMemoryTracking=summary"
# Then, on the host:
PID=$(pgrep -f 'proc_regionserver')
grep VmRSS /proc/$PID/status # what the kernel charges
jcmd $PID GC.heap_info # heap committed and used
jcmd $PID VM.native_memory summary scale=MB # every JVM-known category
jcmd $PID VM.native_memory baseline # snapshot now ...
jcmd $PID VM.native_memory summary.diff scale=MB # ... and compare hours laterNMT's summary lists the heap, class metadata, threads, code, GC and other categories, each with reserved and committed sizes. On recent JDKs, direct ByteBuffer memory appears under the Other category (Internal on JDK 8), so it should agree with the java.nio:type=BufferPool,name=direct MBean. NMT does not see memory that native libraries allocate with malloc themselves. That is the point of the comparison: RSS minus NMT's committed total is roughly what is happening outside the JVM's knowledge.
A worked reconciliation makes the method concrete. The numbers are illustrative. A RegionServer budgeted at 24 GB heap and 72 GB direct reports 101 GB of RSS, three gigabytes more than the expected 98.
| Layer | Meter | Reading | Notes |
|---|---|---|---|
| Heap | GC.heap_info | 24.0 GB | Xms equals Xmx, fully committed |
| Direct | BufferPool MBean / NMT Other | 70.6 GB | cache 60, MSLAB 8, pool and netty 2.6 |
| JVM native | NMT: class, thread, code, GC | 2.9 GB | GC structures 1.4 GB, 900 threads |
| NMT total | sum | 97.5 GB | |
| RSS | /proc status | 101.0 GB | |
| Residue | RSS minus NMT | 3.5 GB | outside the JVM: investigate |
Here the direct layer is inside its limit and the JVM categories are plausible, so the 3.5 GB residue is native allocation the JVM cannot see. The next step is to watch whether it grows. A stable residue is overhead to budget for; a residue that grows every day is a leak.
Reference-count leaks in pooled buffers
HBase's off-heap read path hands out pooled buffers that are reference-counted. A buffer goes back to the ByteBuffAllocator pool when its count reaches zero. If code takes a reference and never releases it, the buffer is eventually garbage-collected along with its wrapper, but it never returns to the pool. The pool drains, allocations fall back to the heap, and the heap allocation ratio that the RegionServer reports climbs. The reverse bug, a release too many, is worse: a buffer is reused while someone still reads it, and clients see wrong values.
The usual sources are custom coprocessors and filters that keep cells beyond the call that produced them, and, bugs in HBase itself; HBASE-27170, fixed in 2.4.14 and 2.5.0, was a pool leak in block decompression. Current HBase source wires the shaded Netty resource leak detector into the RefCnt class, so a pooled buffer collected without being released is logged with the code paths that last touched it. Check that your release has it. The detector is controlled by Netty's leak-detection level property under HBase's relocated thirdparty prefix; confirm the exact name against your build. Its levels trade overhead for detail: simple samples a small fraction of buffers, advanced also records access points, and paranoid tracks every buffer and is only for tests.
# Canary RegionServer only: sampled leak reports with access records.
# Netty reads io.netty.leakDetection.level; HBase shades Netty, so the name below is
# the relocated form. Verify it against your HBase build before relying on it.
export HBASE_REGIONSERVER_OPTS="$HBASE_REGIONSERVER_OPTS \
-Dorg.apache.hbase.thirdparty.io.netty.leakDetection.level=advanced"
# Leak reports are logged at ERROR and start with LEAK:
grep -n 'LEAK:' /var/log/hbase/hbase-*-regionserver-*.log | headRead the reported access records from the bottom up: the last record before the leak usually names the coprocessor hook or filter that took the reference. The fix is almost always the one the reference guide prescribes, which is to copy the bytes you need out of the cell before the hook returns. The coprocessors article covers the hook lifecycle.
Native growth outside the JVM
When the residue outside NMT grows, the JVM is not the allocator. Three causes cover most cases.
- C allocator fragmentation. glibc's malloc creates arenas per thread up to a multiple of the core count, and freed memory in one arena is not returned to the operating system eagerly. A RegionServer with hundreds of threads doing native compression can accumulate gigabytes this way. Capping arenas with
MALLOC_ARENA_MAX(2 or 4 is common) in the RegionServer environment, or switching to jemalloc throughLD_PRELOAD, usually flattens the curve. Test the change on one node, since fewer arenas can raise lock contention. - Native codecs. Compression libraries used through Hadoop's native code allocate their own buffers. These scale with concurrent compactions and flushes, so a residue that rises during compaction storms and falls after is a sign of this, not of a leak.
- Thread growth. Each thread has a native stack. NMT shows thread count in its Thread category; a count that only rises points at a pool without a bound, often in a client library or coprocessor.
To find which mappings grew, compare pmap -x output or /proc/PID/smaps over time. Many anonymous regions of about 64 MB are the signature of glibc arenas.
Rolling out off-heap changes
Moving a production cluster to off-heap caching, or resizing it, is a change to every RegionServer's memory layout, so treat it like a deployment. A safe sequence is:
- Baseline. Record, per RegionServer, a week of p99 GC pause, p99 read and write latency, block cache hit ratio, RSS and direct-memory use.
- Canary. Apply the new
hbase-env.shandhbase-site.xmlto one RegionServer. Drain it first with the region mover (hbase org.apache.hadoop.hbase.util.RegionMover -r HOST -o unload), restart, then load the regions back so the change is invisible to clients. - Warm. An off-heap BucketCache starts empty after a restart. Expect latency to be high until the hit ratio recovers, and compare only after it has.
- Compare. GC pauses should drop and hit ratio should match or exceed the baseline. RSS should settle at the predicted total and stay flat over several days, including a major compaction window.
- Widen. Roll to a rack, then the cluster, one RegionServer at a time with the same drain and load steps. Keep the previous configuration files so rollback is a restart.
Alerts
| Signal | Source | Alert when |
|---|---|---|
| RSS against limit | /proc or container metrics | above 90% of host budget or cgroup limit |
| Direct memory used | BufferPool direct MBean | above 95% of MaxDirectMemorySize |
| Allocator heap allocation ratio | RegionServer metrics | above the reference guide threshold |
| RSS minus NMT committed | canary with NMT | grows across consecutive days |
| Leak reports | RegionServer log | any line starting with LEAK: |
Trend the first two per RegionServer, not as a cluster average; one leaking node is invisible in a mean. The troubleshooting article covers what to collect before restarting a node that is close to its limit.
Failure modes
The failure modes, in the order they tend to appear after a rollout:
- Container kill with no Java error. RSS exceeded the cgroup limit while every JVM meter was inside its own cap. The container was sized for heap and direct only; add JVM native and the observed residue.
- OutOfMemoryError: Direct buffer memory after a config change. Handler count or cache size went up and the direct limit did not. Recompute the direct budget.
- Slow, steady direct growth with a falling pool. A reference-count leak. Run the leak detector on a canary.
- Slow residue growth with flat direct memory. Native fragmentation or codec buffers. Cap arenas and watch through a compaction cycle.
- Latency spike after every restart. Cold BucketCache. Drain and load regions, and budget for warm-up.
Trade-offs
Off-heap memory trades GC pauses for operational work. Measuring it needs NMT, which has a cost, and leak detection, which has a larger one, so both belong on a canary rather than everywhere. Tight budgets use the host well but leave no room for the residue, so keep a few gigabytes of headroom. Capping malloc arenas saves memory and may cost some throughput. A small on-heap cache with a modern collector, described in the GC tuning article, remains a valid choice when the hot set is small.
What to do next
- Confirm that
HBASE_OFFHEAPSIZEis set explicitly on every RegionServer. - Pick one RegionServer, enable NMT in summary mode, and reconcile RSS against heap, direct and JVM native.
- Take an NMT baseline and diff it daily for a week, including a major compaction.
- If direct memory grows while the pool drains, enable advanced leak detection on that node and read the LEAK reports.
- If the residue outside NMT grows, try MALLOC_ARENA_MAX on that node and compare the curve.
- Add the five alerts above, per RegionServer.
- Roll any off-heap change out with drain, restart, warm-up and comparison, one node at a time.