An HBase cluster exposes hundreds of metrics, and most operators either scrape all of them and look at none, or pick a handful by name without knowing what they measure. Both fail at 3 a.m. when p99 latency spikes and the question is which layer is at fault. Useful monitoring starts from how a request flows: into an RPC queue, onto a handler, into a region, through the MemStore and WAL for writes or the block cache and HFiles for reads, down to HDFS, all inside a JVM that can pause.
This article explains how HBase produces metrics, how to get them out, what the important ones mean mechanically with their exact names, and how to read them layer by layer to find a fault. It covers metrics and dashboards; turning them into alert rules and on-call routing is covered in the HBase alerting article.
How HBase produces metrics
HBase uses the Hadoop metrics2 framework. Each subsystem registers a source, and each source appears as a JMX MBean named Hadoop:service=HBase,name=<process>,sub=<source>. On a RegionServer the most important beans are name=RegionServer,sub=Server for server-wide aggregates, sub=IPC for the RPC layer, and sub=Regions and sub=Tables for per-region and per-table breakdowns. On the Master, name=Master,sub=AssignmentManager holds region-in-transition metrics. The JVM's own metrics appear under a JvmMetrics bean.
Metrics come in three shapes. Gauges are current values, such as memStoreSize or regionCount. Counters only increase from process start, such as readRequestCount or updatesBlockedTime, and are meaningless as raw values: compute rates over a window, and expect them to reset on restart. Histograms summarise latency distributions for operations such as get and put, exposed as several fields per operation, including an operation count, mean, and percentiles. The exact field suffixes and capitalisation differ between versions, so read them from your own /jmx output rather than copying names from elsewhere.
A subtlety of histogram percentiles: they are computed per server over a recent window. You cannot average p99 values across servers to get a cluster p99. Plot per-server p99 and look at the worst, or measure latency at the client.
Getting metrics out
Every HBase daemon serves its beans as JSON at /jmx on its web UI port, 16030 for RegionServers and 16010 for the Master by default. The qry parameter filters to one bean, which is much cheaper than fetching everything. From HBase 2.6.0 there is also a /prometheus endpoint serving the same metrics in Prometheus text format; which servlets are enabled is controlled by hbase.http.metrics.servlets. On older versions, the common approach is the Prometheus JMX exporter running as a Java agent.
# One bean from one RegionServer (default info port 16030; Master uses 16010)
curl -s 'http://rs1.example.internal:16030/jmx?qry=Hadoop:service=HBase,name=RegionServer,sub=Server'
# RPC queues and handlers
curl -s 'http://rs1.example.internal:16030/jmx?qry=Hadoop:service=HBase,name=RegionServer,sub=IPC'
# Regions in transition, from the active Master
curl -s 'http://master1.example.internal:16010/jmx?qry=Hadoop:service=HBase,name=Master,sub=AssignmentManager'
# Prometheus text format (HBase 2.6.0 and later)
curl -s 'http://rs1.example.internal:16030/prometheus' | headFor ad hoc diagnosis, a short script that samples one bean twice and computes deltas is often faster than any dashboard, because it forces you to treat counters as rates.
import json
import time
import urllib.parse
import urllib.request
BEAN = "Hadoop:service=HBase,name=RegionServer,sub=Server"
KEYS = ["regionCount", "memStoreSize", "flushQueueLength", "compactionQueueLength",
"storeFileCount", "percentFilesLocal", "updatesBlockedTime", "blockedRequestCount",
"hlogFileCount", "blockCacheHitCount", "blockCacheMissCount", "slowGetCount"]
def snapshot(host, port=16030):
url = f"http://{host}:{port}/jmx?qry=" + urllib.parse.quote(BEAN, safe=":=,")
with urllib.request.urlopen(url, timeout=5) as r:
beans = json.load(r)["beans"]
bean = beans[0] if beans else {}
return {k: bean.get(k) for k in KEYS}
a = snapshot("rs1.example.internal")
time.sleep(60)
b = snapshot("rs1.example.internal")
hits = b["blockCacheHitCount"] - a["blockCacheHitCount"] # counters: use deltas
miss = b["blockCacheMissCount"] - a["blockCacheMissCount"]
print("hit ratio last minute:", hits / max(hits + miss, 1))
print("blocked-update ms last minute:", b["updatesBlockedTime"] - a["updatesBlockedTime"])
print("gauges now:", {k: b[k] for k in ("memStoreSize", "flushQueueLength",
"compactionQueueLength", "hlogFileCount")})If you run the JMX exporter, keep the rule set small and explicit. Exporting every bean, including per-region ones, is the most common way to overload a monitoring system with HBase.
# jmx_exporter rules (Java agent on each RegionServer). Check the exporter's own
# output and adjust patterns: bean property order must match what it reports.
lowercaseOutputName: true
rules:
- pattern: 'Hadoop<service=HBase, name=RegionServer, sub=(Server|IPC)><>(\w+)'
name: hbase_regionserver_$2
labels:
sub: $1
- pattern: 'Hadoop<service=HBase, name=Master, sub=AssignmentManager><>(rit\w+)'
name: hbase_master_$1
The RPC layer: is the server keeping up?
Every client request enters a call queue and waits for a handler thread. queueSize is the bytes of calls waiting; numCallsInGeneralQueue, numCallsInPriorityQueue and numCallsInReplicationQueue count waiting calls per queue type; numActiveHandler and its general, priority and replication variants count busy handlers; numOpenConnections counts client connections.
Read these together. If active handlers equal the configured handler count and the general queue is growing, the server is saturated: every thread is busy and new work waits. The question is then what the handlers are waiting on, which the lower layers answer. If handlers are idle but client latency is high, the problem is outside the server: network, client-side retries or a slow hbase:meta lookup. The priority queue carries meta and admin traffic; if it backs up, the whole cluster slows.
The write path: MemStore, flushes and the WAL
A write appends to the WAL and inserts into the region's MemStore; when a MemStore fills, it is flushed to a new HFile. The metrics follow that sequence. memStoreSize is total MemStore memory on the server. flushQueueLength is how many regions are waiting to flush. hlogFileCount is the number of WAL files; when it exceeds the configured maximum, HBase forces flushes so old WALs can be archived. The write path article explains the mechanics.
The two metrics that mean users are hurting are updatesBlockedTime, a counter of milliseconds during which writes were blocked because a MemStore hit its blocking limit, and blockedRequestCount, the number of requests rejected or held for that reason. Any sustained growth in either means writes are outrunning flushes. The usual chain is: too many store files, so flushes are held back waiting for compaction, so MemStores grow to the blocking size, so writes stop. That is why write-path trouble often shows up first in compactionQueueLength.
slowPutCount, slowAppendCount and slowDeleteCount count operations that exceeded the slow-operation threshold, and put latency histograms show the distribution. Rising slow puts with flat blocked time usually point at WAL sync latency, which means HDFS DataNode or disk trouble.
The read path: cache, files and locality
A get checks the MemStore and then reads blocks from HFiles, first from the block cache and otherwise from HDFS. blockCacheHitCount and blockCacheMissCount are counters; the server also exposes blockCacheCountHitPercent, the hit percentage over all requests, and blockCacheExpressHitPercent, the hit percentage only over requests that asked to be cached. The express ratio is the better signal: scans that disable caching drag down the first without indicating a problem. blockCacheEvictionCount rising fast with a falling hit ratio means the working set no longer fits. The block cache article covers sizing.
storeFileCount is total HFiles on the server. Each read may consult every HFile in a store whose key range and bloom filter do not rule it out, so read latency grows with files per store. That makes compaction a read-path concern: if compactionQueueLength stays high for hours, file counts climb and gets slow down. The compaction article explains the policies.
percentFilesLocal is the percentage of a server's HFile data stored on the local DataNode. After a restart or region move it drops, reads cross the network, and latency rises until major compaction rewrites files locally. A low value after rebalancing explains a latency regression that nothing else does.
The Master: regions in transition
The Master's AssignmentManager bean reports ritCount, the number of regions currently moving between states; ritCountOverThreshold, how many have been in transition longer than a configured threshold; and ritOldestAge, the age in milliseconds of the oldest one. Brief spikes in ritCount during balancing, splits or restarts are normal. A non-zero ritCountOverThreshold or a steadily growing ritOldestAge means a region is stuck, and a stuck region is unavailable to every client that needs it. That is one of the few HBase metrics that should page someone immediately.
Also watch regionCount across RegionServers. A server holding far more regions than average carries more MemStores, more flushes and more requests, and usually shows problems first.
The JVM and the host
Everything above runs inside a JVM, and a long garbage-collection pause stops all of it at once: handlers, flushes and ZooKeeper heartbeats. If a pause outlasts the ZooKeeper session timeout, the Master declares the server dead and reassigns its regions, turning a pause into an outage. Track GC time per interval from the JvmMetrics bean (for example GcTimeMillis, a counter), heap used against maximum, and watch RegionServer logs for the JVM pause monitor's warnings. The GC tuning article covers causes and settings.
At the host level, disk latency on DataNodes, network retransmits and CPU steal on virtual machines explain many slow puts and slow gets that no HBase metric names directly.
Per-region and per-table metrics
The sub=Regions and sub=Tables beans break request counts, store sizes and latencies down per region or table. They are indispensable for finding a hot region, one whose request rate is far above its neighbours because of row-key design. They are also dangerous for your monitoring system: a cluster with 20,000 regions and thirty metrics per region produces 600,000 time series, and region names change on every split. The metrics architecture article explains why that cardinality is costly.
The practical pattern is to export server-level and table-level metrics continuously and query per-region metrics from /jmx on demand when a server looks hot, or export only the top regions by request rate.
Worked example: diagnosing a p99 spike
Clients report that get p99 rose from 8 ms to 120 ms on one table at 14:00. The dashboard shows the rise concentrated on two of twenty RegionServers.
- RPC layer. On both servers, active handlers are at the maximum and the general queue is growing, so the servers are saturated, not the network.
- Read path.
blockCacheExpressHitPercentis steady at 97 percent, butstoreFileCountdoubled since 13:00 andcompactionQueueLengthhas been above 40 for an hour. - Write path. A bulk import started at 13:00.
flushQueueLengthis elevated andupdatesBlockedTimeis growing on the same two servers. - Cause. The import's row keys land in a narrow key range held by regions on those two servers. Flushes create many small HFiles faster than compaction can merge them, so each get consults more files, handlers take longer, and the queue grows.
- Action. Throttle the import, pre-split the key range so load spreads across servers, and let compaction catch up. Confirm recovery by watching store file counts fall and p99 return to baseline.
No single metric showed the cause. The layer-by-layer order did: saturation at the top, files and compaction in the middle, and a workload change underneath.
A dashboard that fits on one screen
| Row | Panels | Question it answers |
|---|---|---|
| Client view | get/put p99 per server, request rate per table | Are users hurting, and where? |
| RPC | general queue length, active handlers, open connections | Is the server saturated? |
| Write path | memStoreSize, flushQueueLength, updatesBlockedTime rate, hlogFileCount | Are writes being blocked? |
| Read path | express hit %, evictions rate, storeFileCount, compactionQueueLength, percentFilesLocal | Why are reads slow? |
| Master | ritCount, ritCountOverThreshold, ritOldestAge, regionCount per server | Is anything unavailable or skewed? |
| JVM and host | GC time rate, heap used, disk latency | Is the process or machine stalling? |
# block cache hit ratio per server over 5 minutes (counters -> rates)
sum by (instance) (rate(hbase_regionserver_blockcachehitcount[5m]))
/ (sum by (instance) (rate(hbase_regionserver_blockcachehitcount[5m]))
+ sum by (instance) (rate(hbase_regionserver_blockcachemisscount[5m])))
# fraction of each minute that writes were blocked (updatesBlockedTime is ms)
rate(hbase_regionserver_updatesblockedtime[1m]) / 1000
# region-count skew across servers
max(hbase_regionserver_regioncount) / avg(hbase_regionserver_regioncount)The queries assume the exporter rules above, with lower-cased names; adapt them to the names your exporter or the /prometheus endpoint actually emits.
Failure modes of the monitoring itself
| Symptom | Cause | Fix |
|---|---|---|
| Counters plotted as flat lines or sawtooths | Raw counters instead of rates | Plot rate() and handle restarts |
| Monitoring backend overloaded | Per-region metrics exported continuously | Export server and table level; query regions on demand |
| Cluster p99 looks fine while users complain | Averaging per-server percentiles | Plot per-server p99 and client-side latency |
| Metrics gap during an incident | Scrape timeout on a fetch of every bean | Scrape specific beans with qry; raise timeout modestly |
| Hit ratio looks poor with no latency impact | Using the all-requests hit percentage | Watch the express hit percentage |
| Dashboards break after upgrade | Metric names or fields changed between versions | Pin names from /jmx per version; test dashboards on upgrade |
What to do next
- Run the curl commands above against one RegionServer and the Master, and record the exact names and histogram fields your version emits.
- Build the six-row dashboard, plotting every counter as a rate.
- Export server-level and table-level beans only; document how to fetch per-region metrics on demand.
- Walk through the worked example against your own cluster during a bulk load or major compaction to learn its normal ranges.
- Add paging alerts for ritCountOverThreshold and sustained updatesBlockedTime, following the alerting article.
- Review the dashboard after every HBase upgrade and fix any renamed metrics before the next incident needs them.