Most HBase metrics are meaningless as raw numbers. A compaction queue of 40, a MemStore of 9 GB or 25 busy handlers might be fine or might be minutes from an outage, and the difference is the limit each one is approaching. HBase has a configured ceiling behind almost every serious incident: a store-file count at which writes block, a MemStore size at which a region refuses updates, a call-queue length at which requests are rejected, a WAL-file count at which flushes are forced, and a pause length at which a RegionServer loses its ZooKeeper session.

This article pairs each important metric with the limit it runs into and with what HBase does when it gets there, then turns the pairs into a headroom calculation you can run, trend and alert on. It assumes you can already collect metrics; how the metrics2 pipeline, JMX beans and endpoints work is in HBase metrics and monitoring, and how to aggregate counters, gauges and histograms correctly is in HBase metrics deep dive. Metric names and defaults here were read from the HBase 2.6 branch.

Metrics and their limits in one picture

MetricLimit (default)At the limitstoreFileCountper storeblockingStoreFiles16Region updates blockedup to 90 sregion memStoreSizeper regionflush.size x multiplier128 MB x 4RegionTooBusyExceptionwrites rejectedserver memStoreSizesub=Serverglobal memstore0.4 of heapForced flushes, then blocklower limit 95%numCallsInGeneralQueuesub=IPC10 x handlers per queuehandlers 30callQueueTooBigclient retrieshlogFileCountlive WAL filesmaxlogsmax(32, computed)Oldest regions flushedflush stormJVM pauseGC or hostzookeeper.session.timeout90 s, capped by ensembleSession expiresRegionServer abortsheadroom = 1 - value / limit, tracked per server, alerted when it trends to zero
Each metric approaches a configured limit, and at the limit HBase changes behaviour: blocking updates, rejecting calls, forcing flushes or aborting. Headroom measures the distance.

Why headroom beats raw values

Headroom is the fraction of a limit still unused: one minus the value divided by the limit. It has three advantages over raw values. It is comparable across metrics, so one dashboard row can show whether the worst constraint on each server is store files or the call queue. It survives configuration changes, because when you raise hbase.hstore.blockingStoreFiles the headroom rises with it, whereas a hard-coded alert at 12 store files would silently become wrong. And it can be forecast: a headroom series that falls by a steady amount each hour tells you when it will reach zero.

The catch is that you must know the limit, which means reading configuration and heap size, not only metrics. Limits can also be overridden per table in the table descriptor, so a cluster-wide value is a starting point, not the truth for every region.

The limit table

The pairs below cover most write stalls, rejections and RegionServer losses. Bean names are under Hadoop:service=HBase,name=RegionServer unless marked as Master.

Metric (bean)Limit and defaultWhat happens at the limitEvidence metric
storeFileCount (sub=Server is a server total; per-region in sub=Regions)hbase.hstore.blockingStoreFiles = 16 per storeUpdates to the region block until compaction catches up, for up to hbase.hstore.blockingWaitTime = 90,000 msupdatesBlockedTime, blockedRequestCount
Per-region memStoreSize (sub=Regions)hbase.hregion.memstore.flush.size 128 MB x hbase.hregion.memstore.block.multiplier 4 = 512 MBThe region rejects writes with RegionTooBusyExceptionexceptions.RegionTooBusyException
memStoreSize (sub=Server)hbase.regionserver.global.memstore.size = 0.4 of heap; lower limit 0.95 of thatAbove the lower limit, flushes are forced; above the upper limit, all updates blockupdatesBlockedTime
numCallsInGeneralQueue, queueSize (sub=IPC)10 calls per handler serving the queue; hbase.ipc.server.max.callqueue.size = 1 GBNew calls rejected; clients back off and retryexceptions.callQueueTooBig
numActiveHandler (sub=IPC)hbase.regionserver.handler.count = 30Calls wait in the queue; queue time growsqueueCallTime
hlogFileCounthbase.regionserver.maxlogs = max(32, 2 x global memstore / WAL roll size)Regions holding the oldest WAL are flushed so it can be archivedflushQueueLength
ritCountOverThreshold (Master, sub=AssignmentManager)hbase.metrics.rit.stuck.warning.threshold = 60,000 msCounts regions in transition longer than the threshold; their data is unavailableritOldestAge (longest current RIT, ms)
JVM pause lengthzookeeper.session.timeout = 90,000 ms requested; an external ensemble's maxSessionTimeout caps it, and ZooKeeper ships with 40 sThe session expires and the RegionServer aborts; its regions are reassignedGC time; pause warnings in the log

Write-path headroom

Three write-path limits interact, and the order in which they bite tells you what is wrong. When compaction falls behind, store files accumulate; compactionQueueLength has no limit of its own, but a growing queue is the leading indicator, and store-file headroom is the trailing one. When a region's store reaches 16 files, its writes block; updatesBlockedTime, a counter, starts increasing, and its rate is the most direct measure of write stalls users feel. Compaction tuning is covered in HBase compaction.

MemStore pressure comes from two directions. One hot region can hit its own 512 MB block limit while the server has plenty of memory, and the server-wide metric will look healthy; only per-region MemStore sizes or RegionTooBusyException counts show it. Alternatively, many regions together can push the server past its global lower limit, and HBase starts flushing the largest MemStores regardless of their own size, creating many small files that then feed compaction. The flush mechanics are explained in MemStore flush.

WAL files close the loop. A WAL file can be archived only when every region with edits in it has flushed those edits. A rarely written region can pin old WAL files for hours, and when the live count exceeds hbase.regionserver.maxlogs, HBase forces flushes of the regions holding the oldest file. Watch hlogFileCount against the limit and the slowAppendCount and syncTime metrics in the WAL source; the WAL logs a slow sync above hbase.regionserver.wal.slowsync.ms, default 100 ms.

RPC headroom

Every request waits in a call queue until a handler thread is free. With the default 30 handlers, numActiveHandler at 30 means every handler is busy, and from then on latency grows with queueCallTime, not with processing time. The queue's length limit is 10 calls per handler serving that queue unless hbase.ipc.server.max.callqueue.length is set, and its byte limit is 1 GB. Beyond either, new calls are rejected and counted in exceptions.callQueueTooBig. Do not confuse this with numGeneralCallsDropped, which counts calls dropped by the CoDel queue policy when it is enabled.

High handler utilisation with short process times means too few handlers for the request rate. High utilisation with long process times means handlers are waiting on something below, usually a blocked write path or slow HDFS, and adding handlers only deepens the queue in front of the real bottleneck.

Liveness: pauses and regions in transition

Two limits decide whether a server survives at all. A JVM pause, from garbage collection or from the host stalling the process, that outlasts the ZooKeeper session timeout loses the RegionServer's session, and the server aborts. HBase requests zookeeper.session.timeout, 90 seconds by default, but an ensemble HBase does not manage applies its own maxSessionTimeout when lower, and ZooKeeper ships with 40 seconds, so read the negotiated value from the RegionServer log or the ensemble. Track the longest pause per server against that effective timeout; a server whose pauses reach a few seconds has far less headroom than its average GC time suggests. On the Master, ritCount should return to zero after any balancing or failover, and ritCountOverThreshold above zero means some regions have been in transition for more than a minute and are unavailable to clients.

A headroom calculator

The script reads the server and IPC beans from a RegionServer's /jmx endpoint and computes headroom against limits derived from configuration you pass in. It deliberately takes the limits as inputs: the right values depend on your hbase-site.xml, your heap, and table-level overrides, and a script that guessed them would be wrong silently.

import json
import urllib.parse
import urllib.request

def bean(host, port, qry):
    url = f"http://{host}:{port}/jmx?qry=" + urllib.parse.quote(qry, safe=":=,")
    with urllib.request.urlopen(url, timeout=5) as r:
        beans = json.load(r)["beans"]
    return beans[0] if beans else {}

def limits(heap_bytes, handlers=30, roll_bytes=128 * 2**20):
    """Limits computed from config you supply; check each against your hbase-site.xml."""
    global_upper = 0.4 * heap_bytes
    return {
        "memStoreSize": 0.95 * global_upper,     # forced flushing starts here
        "hlogFileCount": max(32, int(2 * global_upper / roll_bytes)),
        "numActiveHandler": handlers,
        # Total across general queues; with several queues, one imbalanced
        # queue can reject calls before this total is reached.
        "numCallsInGeneralQueue": 10 * handlers,
    }

def headroom(host, heap_bytes, **cfg):
    server = bean(host, 16030, "Hadoop:service=HBase,name=RegionServer,sub=Server")
    ipc = bean(host, 16030, "Hadoop:service=HBase,name=RegionServer,sub=IPC")
    values = {**server, **ipc}
    return {k: round(1 - values.get(k, 0) / lim, 3)
            for k, lim in limits(heap_bytes, **cfg).items()}

print(headroom("rs1.example.internal", heap_bytes=31 * 2**30))

Extend it with per-region beans from sub=Regions, whose attribute names follow the pattern Namespace_<ns>_table_<table>_region_<encoded>_metric_<name>, to compute store-file and MemStore headroom for the worst region rather than the server total. Divide a region's store-file count by its number of column families as a rough per-store figure, since stores are per family.

Worked example: a 31 GB RegionServer

Take a RegionServer with a 31 GB heap, the default 30 handlers, and, as an assumption for this example, a WAL roll size of 128 MB; read your own from the logs or configuration. The global MemStore upper limit is 0.4 x 31 GB, about 12.4 GB, and forced flushing starts at 95 percent of that, about 11.8 GB. The WAL-file limit is the larger of 32 and 2 x 12.4 GB / 128 MB, which is 198 files.

One afternoon the server reports memStoreSize at 10.9 GB, hlogFileCount at 171, numActiveHandler at 12, and its worst region has 14 store files in a single-family table. Headroom is 8 percent on global MemStore, 14 percent on WAL files, 60 percent on handlers and 12.5 percent on store files. The handlers are fine; the write path is not. Store files are two short of blocking, and the WAL count is near the point where forced flushes will create more small files, which will push store files over. The order of the fix follows: find the rarely written regions pinning old WAL files and flush them, check that compactionQueueLength is draining rather than growing, and only then consider raising limits.

Expressed as recording and alert rules, a headroom series makes this visible hours earlier. The 198 below is this example's computed limit, not a default.

# Recording rules: headroom from gauges you already scrape. Metric names follow
# whatever your exporter produces; adjust them to your own scrape output.
- record: hbase:handler_headroom
  expr: 1 - hbase_regionserver_numActiveHandler / 30
- record: hbase:wal_headroom
  expr: 1 - hbase_regionserver_hlogFileCount / 198
# Page when the worst headroom is forecast to reach zero within 30 minutes.
- alert: HBaseWalHeadroomExhausting
  expr: predict_linear(hbase:wal_headroom[1h], 1800) < 0
  for: 10m

How these rules fit into routing and on-call is covered in HBase alerting.

Failure modes

  • Server totals hide the worst region. A server-level store-file count of 900 says nothing about the one store at 15. Compute headroom per region for the write-path limits.
  • Wrong limits. Table-level overrides, a changed heap, or off-heap MemStore change the ceiling. Recompute limits whenever configuration or heap changes, and record them alongside the metrics.
  • Counters treated as gauges. updatesBlockedTime and the exception counts only grow; alert on their rate, not their value.
  • Averages across servers. One server at zero headroom matters more than a healthy average. Alert on the minimum.
  • Headroom gamed by raising limits. Raising hbase.hstore.blockingStoreFiles restores headroom on paper while read amplification and compaction debt grow. Treat a raised limit as a change to review, not a fix.
  • Missing liveness data. A RegionServer that stopped reporting has no headroom at all; alert on absent series as well as low values.

Trade-offs

Headroom monitoring costs more upfront than threshold alerts: you must model limits from configuration, keep that model current, and accept per-region metrics with their cardinality cost. In exchange, alerts survive tuning changes, dashboards show which constraint is closest on each server, and forecasts give warning before a stall rather than after. Raising a limit is a valid response when hardware has room, but it moves the problem to a less visible place, so prefer fixing the rate that consumes the headroom.

What to do next

  1. Write down, for your cluster, the effective value of each limit in the table, including table-level overrides and the WAL roll size.
  2. Compute headroom per server for global MemStore, WAL files, handlers and call queue, using the calculator above.
  3. Add per-region headroom for store files and MemStore on the tables that take most writes.
  4. Record headroom as series and alert on the minimum across servers and on forecasts to zero.
  5. Alert on the rate of updatesBlockedTime, exceptions.RegionTooBusyException and exceptions.callQueueTooBig as confirmation that a limit was hit.
  6. Track the longest JVM pause per server against the negotiated ZooKeeper session timeout, and ritCountOverThreshold on the Master.
  7. Recompute limits on every configuration or heap change, and review any raised limit for the cost it moves elsewhere.
Key takeaway: An HBase metric means something only next to the limit it approaches. Pair store files with blockingStoreFiles, region MemStore with flush size times the block multiplier, server MemStore with the global limit, handlers and call queues with their caps, WAL files with maxlogs and pauses with the negotiated ZooKeeper session timeout. Compute headroom from your real configuration, track the worst region and server, forecast it to zero, and fix the rate that consumes it before raising the limit.