Most HBase metrics are meaningless as raw numbers. A compaction queue of 40, a MemStore of 9 GB or 25 busy handlers might be fine or might be minutes from an outage, and the difference is the limit each one is approaching. HBase has a configured ceiling behind almost every serious incident: a store-file count at which writes block, a MemStore size at which a region refuses updates, a call-queue length at which requests are rejected, a WAL-file count at which flushes are forced, and a pause length at which a RegionServer loses its ZooKeeper session.
This article pairs each important metric with the limit it runs into and with what HBase does when it gets there, then turns the pairs into a headroom calculation you can run, trend and alert on. It assumes you can already collect metrics; how the metrics2 pipeline, JMX beans and endpoints work is in HBase metrics and monitoring, and how to aggregate counters, gauges and histograms correctly is in HBase metrics deep dive. Metric names and defaults here were read from the HBase 2.6 branch.
Metrics and their limits in one picture
Why headroom beats raw values
Headroom is the fraction of a limit still unused: one minus the value divided by the limit. It has three advantages over raw values. It is comparable across metrics, so one dashboard row can show whether the worst constraint on each server is store files or the call queue. It survives configuration changes, because when you raise hbase.hstore.blockingStoreFiles the headroom rises with it, whereas a hard-coded alert at 12 store files would silently become wrong. And it can be forecast: a headroom series that falls by a steady amount each hour tells you when it will reach zero.
The catch is that you must know the limit, which means reading configuration and heap size, not only metrics. Limits can also be overridden per table in the table descriptor, so a cluster-wide value is a starting point, not the truth for every region.
The limit table
The pairs below cover most write stalls, rejections and RegionServer losses. Bean names are under Hadoop:service=HBase,name=RegionServer unless marked as Master.
| Metric (bean) | Limit and default | What happens at the limit | Evidence metric |
|---|---|---|---|
storeFileCount (sub=Server is a server total; per-region in sub=Regions) | hbase.hstore.blockingStoreFiles = 16 per store | Updates to the region block until compaction catches up, for up to hbase.hstore.blockingWaitTime = 90,000 ms | updatesBlockedTime, blockedRequestCount |
Per-region memStoreSize (sub=Regions) | hbase.hregion.memstore.flush.size 128 MB x hbase.hregion.memstore.block.multiplier 4 = 512 MB | The region rejects writes with RegionTooBusyException | exceptions.RegionTooBusyException |
memStoreSize (sub=Server) | hbase.regionserver.global.memstore.size = 0.4 of heap; lower limit 0.95 of that | Above the lower limit, flushes are forced; above the upper limit, all updates block | updatesBlockedTime |
numCallsInGeneralQueue, queueSize (sub=IPC) | 10 calls per handler serving the queue; hbase.ipc.server.max.callqueue.size = 1 GB | New calls rejected; clients back off and retry | exceptions.callQueueTooBig |
numActiveHandler (sub=IPC) | hbase.regionserver.handler.count = 30 | Calls wait in the queue; queue time grows | queueCallTime |
hlogFileCount | hbase.regionserver.maxlogs = max(32, 2 x global memstore / WAL roll size) | Regions holding the oldest WAL are flushed so it can be archived | flushQueueLength |
ritCountOverThreshold (Master, sub=AssignmentManager) | hbase.metrics.rit.stuck.warning.threshold = 60,000 ms | Counts regions in transition longer than the threshold; their data is unavailable | ritOldestAge (longest current RIT, ms) |
| JVM pause length | zookeeper.session.timeout = 90,000 ms requested; an external ensemble's maxSessionTimeout caps it, and ZooKeeper ships with 40 s | The session expires and the RegionServer aborts; its regions are reassigned | GC time; pause warnings in the log |
Write-path headroom
Three write-path limits interact, and the order in which they bite tells you what is wrong. When compaction falls behind, store files accumulate; compactionQueueLength has no limit of its own, but a growing queue is the leading indicator, and store-file headroom is the trailing one. When a region's store reaches 16 files, its writes block; updatesBlockedTime, a counter, starts increasing, and its rate is the most direct measure of write stalls users feel. Compaction tuning is covered in HBase compaction.
MemStore pressure comes from two directions. One hot region can hit its own 512 MB block limit while the server has plenty of memory, and the server-wide metric will look healthy; only per-region MemStore sizes or RegionTooBusyException counts show it. Alternatively, many regions together can push the server past its global lower limit, and HBase starts flushing the largest MemStores regardless of their own size, creating many small files that then feed compaction. The flush mechanics are explained in MemStore flush.
WAL files close the loop. A WAL file can be archived only when every region with edits in it has flushed those edits. A rarely written region can pin old WAL files for hours, and when the live count exceeds hbase.regionserver.maxlogs, HBase forces flushes of the regions holding the oldest file. Watch hlogFileCount against the limit and the slowAppendCount and syncTime metrics in the WAL source; the WAL logs a slow sync above hbase.regionserver.wal.slowsync.ms, default 100 ms.
RPC headroom
Every request waits in a call queue until a handler thread is free. With the default 30 handlers, numActiveHandler at 30 means every handler is busy, and from then on latency grows with queueCallTime, not with processing time. The queue's length limit is 10 calls per handler serving that queue unless hbase.ipc.server.max.callqueue.length is set, and its byte limit is 1 GB. Beyond either, new calls are rejected and counted in exceptions.callQueueTooBig. Do not confuse this with numGeneralCallsDropped, which counts calls dropped by the CoDel queue policy when it is enabled.
High handler utilisation with short process times means too few handlers for the request rate. High utilisation with long process times means handlers are waiting on something below, usually a blocked write path or slow HDFS, and adding handlers only deepens the queue in front of the real bottleneck.
Liveness: pauses and regions in transition
Two limits decide whether a server survives at all. A JVM pause, from garbage collection or from the host stalling the process, that outlasts the ZooKeeper session timeout loses the RegionServer's session, and the server aborts. HBase requests zookeeper.session.timeout, 90 seconds by default, but an ensemble HBase does not manage applies its own maxSessionTimeout when lower, and ZooKeeper ships with 40 seconds, so read the negotiated value from the RegionServer log or the ensemble. Track the longest pause per server against that effective timeout; a server whose pauses reach a few seconds has far less headroom than its average GC time suggests. On the Master, ritCount should return to zero after any balancing or failover, and ritCountOverThreshold above zero means some regions have been in transition for more than a minute and are unavailable to clients.
A headroom calculator
The script reads the server and IPC beans from a RegionServer's /jmx endpoint and computes headroom against limits derived from configuration you pass in. It deliberately takes the limits as inputs: the right values depend on your hbase-site.xml, your heap, and table-level overrides, and a script that guessed them would be wrong silently.
import json
import urllib.parse
import urllib.request
def bean(host, port, qry):
url = f"http://{host}:{port}/jmx?qry=" + urllib.parse.quote(qry, safe=":=,")
with urllib.request.urlopen(url, timeout=5) as r:
beans = json.load(r)["beans"]
return beans[0] if beans else {}
def limits(heap_bytes, handlers=30, roll_bytes=128 * 2**20):
"""Limits computed from config you supply; check each against your hbase-site.xml."""
global_upper = 0.4 * heap_bytes
return {
"memStoreSize": 0.95 * global_upper, # forced flushing starts here
"hlogFileCount": max(32, int(2 * global_upper / roll_bytes)),
"numActiveHandler": handlers,
# Total across general queues; with several queues, one imbalanced
# queue can reject calls before this total is reached.
"numCallsInGeneralQueue": 10 * handlers,
}
def headroom(host, heap_bytes, **cfg):
server = bean(host, 16030, "Hadoop:service=HBase,name=RegionServer,sub=Server")
ipc = bean(host, 16030, "Hadoop:service=HBase,name=RegionServer,sub=IPC")
values = {**server, **ipc}
return {k: round(1 - values.get(k, 0) / lim, 3)
for k, lim in limits(heap_bytes, **cfg).items()}
print(headroom("rs1.example.internal", heap_bytes=31 * 2**30))Extend it with per-region beans from sub=Regions, whose attribute names follow the pattern Namespace_<ns>_table_<table>_region_<encoded>_metric_<name>, to compute store-file and MemStore headroom for the worst region rather than the server total. Divide a region's store-file count by its number of column families as a rough per-store figure, since stores are per family.
Worked example: a 31 GB RegionServer
Take a RegionServer with a 31 GB heap, the default 30 handlers, and, as an assumption for this example, a WAL roll size of 128 MB; read your own from the logs or configuration. The global MemStore upper limit is 0.4 x 31 GB, about 12.4 GB, and forced flushing starts at 95 percent of that, about 11.8 GB. The WAL-file limit is the larger of 32 and 2 x 12.4 GB / 128 MB, which is 198 files.
One afternoon the server reports memStoreSize at 10.9 GB, hlogFileCount at 171, numActiveHandler at 12, and its worst region has 14 store files in a single-family table. Headroom is 8 percent on global MemStore, 14 percent on WAL files, 60 percent on handlers and 12.5 percent on store files. The handlers are fine; the write path is not. Store files are two short of blocking, and the WAL count is near the point where forced flushes will create more small files, which will push store files over. The order of the fix follows: find the rarely written regions pinning old WAL files and flush them, check that compactionQueueLength is draining rather than growing, and only then consider raising limits.
Expressed as recording and alert rules, a headroom series makes this visible hours earlier. The 198 below is this example's computed limit, not a default.
# Recording rules: headroom from gauges you already scrape. Metric names follow
# whatever your exporter produces; adjust them to your own scrape output.
- record: hbase:handler_headroom
expr: 1 - hbase_regionserver_numActiveHandler / 30
- record: hbase:wal_headroom
expr: 1 - hbase_regionserver_hlogFileCount / 198
# Page when the worst headroom is forecast to reach zero within 30 minutes.
- alert: HBaseWalHeadroomExhausting
expr: predict_linear(hbase:wal_headroom[1h], 1800) < 0
for: 10mHow these rules fit into routing and on-call is covered in HBase alerting.
Failure modes
- Server totals hide the worst region. A server-level store-file count of 900 says nothing about the one store at 15. Compute headroom per region for the write-path limits.
- Wrong limits. Table-level overrides, a changed heap, or off-heap MemStore change the ceiling. Recompute limits whenever configuration or heap changes, and record them alongside the metrics.
- Counters treated as gauges.
updatesBlockedTimeand the exception counts only grow; alert on their rate, not their value. - Averages across servers. One server at zero headroom matters more than a healthy average. Alert on the minimum.
- Headroom gamed by raising limits. Raising
hbase.hstore.blockingStoreFilesrestores headroom on paper while read amplification and compaction debt grow. Treat a raised limit as a change to review, not a fix. - Missing liveness data. A RegionServer that stopped reporting has no headroom at all; alert on absent series as well as low values.
Trade-offs
Headroom monitoring costs more upfront than threshold alerts: you must model limits from configuration, keep that model current, and accept per-region metrics with their cardinality cost. In exchange, alerts survive tuning changes, dashboards show which constraint is closest on each server, and forecasts give warning before a stall rather than after. Raising a limit is a valid response when hardware has room, but it moves the problem to a less visible place, so prefer fixing the rate that consumes the headroom.
What to do next
- Write down, for your cluster, the effective value of each limit in the table, including table-level overrides and the WAL roll size.
- Compute headroom per server for global MemStore, WAL files, handlers and call queue, using the calculator above.
- Add per-region headroom for store files and MemStore on the tables that take most writes.
- Record headroom as series and alert on the minimum across servers and on forecasts to zero.
- Alert on the rate of
updatesBlockedTime,exceptions.RegionTooBusyExceptionandexceptions.callQueueTooBigas confirmation that a limit was hit. - Track the longest JVM pause per server against the negotiated ZooKeeper session timeout, and
ritCountOverThresholdon the Master. - Recompute limits on every configuration or heap change, and review any raised limit for the cost it moves elsewhere.