HBase problems rarely arrive labelled. What arrives is a page saying writes are timing out, a Spark job stuck at 99 percent, or a dashboard showing 40 regions in transition. The components involved, the Master, RegionServers, ZooKeeper, HDFS and the hbase:meta table, are tightly coupled, so one fault shows up as symptoms in several places, and the most visible symptom is often not the cause.
This article is a set of runbooks organised by symptom. It assumes you already have metrics, which are covered in HBase metrics and monitoring, and focuses on what to do when something is wrong: how to narrow the problem down, which evidence to capture before it disappears, the common failure patterns and their log signatures, how to use HBCK2 without causing damage, and a worked incident from crash to recovery.
Triage: narrow the blast radius first
Before looking at any single error, answer one question: how much of the cluster is affected? The answer decides where to look.
- Every table, every client: suspect something shared:
hbase:metaunavailable, the active Master down, a ZooKeeper quorum problem or HDFS trouble (NameNode, too few DataNodes, full disks). - Everything hosted on one RegionServer: suspect that process or its host: a long GC pause, a slow disk under the WAL, a network problem or an overloaded handler pool.
- One table or a few regions: suspect those regions: stuck in transition, blocked on too many store files, a hot row key or a missing file.
- One client application: suspect the client: stale configuration, a scan without limits, retries that multiply load, or a client-side thread pool that is too small.
Client exceptions help here. RegionTooBusyException points at one region's memstore; CallQueueTooBigException at one server's RPC queue; NotServingRegionException repeating for the same region points at assignment; timeouts on every call point at something shared.
Evidence to collect in the first five minutes
Much of the useful evidence is short-lived: a thread dump, the procedure list while it is stuck, the GC log around a pause. Capture it before restarting anything, because a restart wipes it and often moves the problem somewhere harder to see.
# hbase shell: the first five minutes
status 'simple' # live and dead RegionServers, load per server
status 'detailed' # per-region counts and sizes (large output)
list_procedures # what the Master is doing, and what is stuck
list_locks # which procedure holds which table/region lock
list_regions 'orders' # per-region size, requests and locality
# Is meta reachable, and where is it?
scan 'hbase:meta', {FILTER => "PrefixFilter('orders,')", COLUMNS => ['info:regioninfo', 'info:server', 'info:state'], LIMIT => 5}Also save the Master UI's regions-in-transition and procedures pages, a jstack of any stuck RegionServer, and the last hour of relevant logs. Then search for the lines that explain most incidents:
# On a RegionServer: the lines that explain most incidents
grep -E "Detected pause in JVM or host machine" hbase-*-regionserver-*.log | tail
grep -E "Slow sync cost" hbase-*-regionserver-*.log | tail
grep -E "Blocking updates|Over memstore limit|RegionTooBusyException" hbase-*-regionserver-*.log | tail
grep -E "responseTooSlow|responseTooLarge" hbase-*-regionserver-*.log | tail
grep -E "Failed open of region|ABORTING region server|YouAreDeadException" hbase-*-regionserver-*.log | tail
RegionServer aborts: pauses and expired sessions
A RegionServer proves it is alive through its ZooKeeper session. If the process stops responding for longer than the session timeout (zookeeper.session.timeout, 90 seconds by default, but capped by the ZooKeeper servers' own maximum), the session expires, the Master declares the server dead and starts reassigning its regions. When the paused process wakes up, it learns it has been declared dead and aborts, often with YouAreDeadException in the log.
The cause is almost always a long stop-the-world pause. HBase logs Detected pause in JVM or host machine with the length; if the GC log shows a matching full collection, the heap or collector needs attention, as described in HBase GC tuning. If the GC log shows nothing, the host paused the JVM: swapping, a frozen virtual machine, or a kernel stall on a bad disk. Check vmstat and the kernel log for the same minute.
Do not "fix" this by raising the session timeout far above the typical pause. A longer timeout means a truly dead server holds its regions unavailable for longer. Fix the pause, and keep the timeout short enough that genuine failures recover quickly.
Regions stuck in transition
In HBase 2, every region move, open, close, split and merge is a procedure run by the Master's procedure framework, and a region is "in transition" while its procedure is running. A few regions in transition during a rolling restart is normal. Regions that stay in transition for many minutes mean the procedure cannot complete, and the cause is almost always on the RegionServer that tried to open the region.
Look for Failed open of region in that server's log. The usual reasons are an HFile that is missing or corrupt, an HDFS problem reading the region directory, a table whose compression codec is not installed on that host, or an unreachable hbase:meta. The Master keeps retrying, so the region moves between servers and fails on each one until the cause is fixed. A coprocessor that cannot load behaves differently: with hbase.coprocessor.abortonerror at its default of true, the server aborts, so the symptom is RegionServers dying one after another as the region moves.
Fix the cause first, then let the retry succeed or trigger it with HBCK2. HBCK2 is the operator tool for HBase 2; the old hbck repair flags, which edited meta and HDFS directly, are not supported on HBase 2 and must not be used. Start with the read-only reports, and use bypass only for procedures that can never finish, because it releases the procedure's locks without doing its work and can leave state you then have to repair by hand.
# HBCK2 is a separate jar; the old hbck repair options do not work on HBase 2.x.
HBCK2="hbase hbck -j /opt/hbase-operator-tools/hbase-hbck2.jar"
$HBCK2 reportMissingRegionsInMeta orders # read-only checks first
$HBCK2 extraRegionsInMeta orders
# Region stuck because its last open failed (after fixing the CAUSE):
$HBCK2 assigns d1f3a8c47e9b20c65a4f3e8b7c2d1e09 # encoded region name
# Procedure that can never finish (e.g. its region no longer exists):
$HBCK2 bypass 4711 # read `bypass --help` first; may need override flags
# Master believes a crashed server's WALs were handled when they were not:
$HBCK2 scheduleRecoveries rs7.example.com,16020,1727650000000If the region in question is meta itself, everything else waits: no client can find any region until meta is open. Meta's structure and recovery are covered in the hbase:meta article.
Writes blocked: RegionTooBusyException and store-file pressure
HBase protects itself by blocking writes before memory or read cost gets out of hand. There are two per-region brakes. When a region's memstore reaches hbase.hregion.memstore.block.multiplier (default 4) times the flush size (default 128 MB), that is 512 MB, updates to that region are blocked and clients get RegionTooBusyException until a flush catches up. When a store has more than hbase.hstore.blockingStoreFiles (default 16) files, flushes for that region are delayed until compaction reduces the count or hbase.hstore.blockingWaitTime (default 90 seconds) passes, which in turn pushes the memstore towards the first limit.
The log shows Blocking updates or Over memstore limit. The typical cause is a write burst, such as a backfill or a bulk import through the normal write path, that produces flushes faster than compactions can merge them, often concentrated on a few regions. There is also a server-wide limit on total memstore size across all regions, which blocks every region on the server at once.
Short-term relief: slow the writer (clients should back off, not retry harder), and check that compactions are running and not starved of threads or disk throughput. Longer term: spread the write load across more regions, use bulk loading for backfills instead of millions of puts, and raise the blocking store-file count only if compaction throughput can keep up. Raising limits without fixing throughput only moves the stall to a bigger, later one.
Slow reads, slow calls and the online slow log
RegionServers log calls that exceed a time or size threshold as responseTooSlow or responseTooLarge, with the method, the processing and queue time, and the client. From HBase 2.3 onwards you can also keep recent slow calls in an in-memory ring buffer on each RegionServer and read them from the shell, which is quicker than collecting logs from many servers:
# hbase-site.xml on RegionServers (off by default)
# hbase.regionserver.slowlog.buffer.enabled = true
# hbase.regionserver.slowlog.ringbuffer.size = 256 (the default)
# hbase shell: recent slow calls, then only those for one table
get_slowlog_responses '*'
get_slowlog_responses {'TABLE_NAME' => 'orders'}Separate queue time from processing time. Long queue time with short processing means the handler pool is saturated (hbase.regionserver.handler.count defaults to 30), often by a few expensive scans starving cheap gets. Long processing time for gets usually means many HFiles to merge (compaction behind), poor block-cache hit rates, or low locality after regions moved, so reads go over the network to remote DataNodes. Long processing for scans usually means a filter scanning far more rows than it returns; the fix is a better row key or a tighter start and stop row, not a bigger server.
Hotspots, meta and the WAL
Hotspots. If one RegionServer is overloaded while others are idle, check per-region request counts. Monotonically increasing row keys, such as timestamps or sequence numbers, send every write to the last region. Moving the region only moves the heat; the design fixes are in hotspotting.
Meta. Clients cache region locations, so a brief meta outage is often invisible until caches go stale after a large reassignment. A burst of client lookups at the same moment, for example after a RegionServer crash, can then overload the one server hosting meta. Keep meta's server lightly loaded and watch its request rate during incidents.
The WAL. Every write waits for its WAL sync to HDFS. Slow sync cost lines mean the HDFS write pipeline is slow: a failing disk, a slow DataNode or network congestion. Because writes on that server all share the WAL, one slow DataNode in the pipeline slows every write on the server, which looks like a RegionServer problem but is an HDFS one. After a crash, the dead server's WALs must be split and replayed before its regions open, so slow WAL splitting shows up as a long recovery time. The mechanics of the RegionServer process that ties these together are in the RegionServer article.
Worked incident: a crash, a replay and one region that will not open
02:10. Alerts fire: write latency on the orders table is up, and one RegionServer, rs7, is missing. The client errors are NotServingRegionException and timeouts, and only regions previously on rs7 are affected, so this is a one-server problem.
02:12. rs7's log ends with Detected pause in JVM or host machine reporting about 95 seconds, followed by YouAreDeadException and an abort. Its GC log shows a full collection during a large scan. The session expired, the Master split rs7's WALs and started reassigning its 180 regions to other servers. By 02:16, 179 are open.
02:17. One region of orders is still in transition, and two of the three RegionServers added last month have just aborted. Their logs show an audit coprocessor failing to load with ClassNotFoundException: its jar went to rs1 to rs7 but never to the new servers. With hbase.coprocessor.abortonerror at its default of true, each new server that received the region aborted, and because the new servers were the least loaded, the Master kept choosing them.
02:25. The team copies the jar to the new servers and restarts the two aborted RegionServers. The region opens on the next retry (had its procedure given up, HBCK2 assigns would restart it) and write latency returns to normal. Nobody restarted the Master, deleted ZooKeeper nodes or used bypass, because none of those would have fixed a missing jar.
Follow-up: fix the scan behind the pause, bake coprocessor jars into the server image, and alert on regions in transition for over five minutes.
Fixes that make things worse
| Tempting action | Why it backfires | Do instead |
|---|---|---|
| Restarting the Master repeatedly | Each restart reloads procedures; the stuck one resumes and fails for the same reason | Find the failing open in the RegionServer log |
| Deleting ZooKeeper nodes | HBase 2 keeps assignment state in meta and the procedure store; you remove clues, not the cause | Leave ZooKeeper alone unless a documented procedure says otherwise |
| Old hbck repair flags | Not supported on HBase 2; can corrupt meta | HBCK2, read-only reports first |
bypass as a first resort | Releases locks without doing the work | Only for procedures that can never complete |
| Raising every timeout and limit | Hides the fault and slows failure recovery | Fix pauses, compaction throughput and hot keys |
| Major compaction of everything during the incident | Adds I/O to a struggling cluster | Compact specific regions after recovery |
What to do next
- Write the triage question (cluster, server, table or client?) at the top of your on-call runbook.
- Script the first-five-minutes evidence collection: shell status, procedures, locks, thread dumps and log greps.
- Install HBCK2 on an edge node now, matching your HBase version, and practise its read-only reports on staging.
- Enable the online slow log on RegionServers and add slow-call queue versus processing time to your dashboards.
- Alert on regions in transition for more than five minutes, JVM pauses over 10 seconds, and blocked updates.
- Check that every RegionServer, including recently added ones, has the same coprocessors, codecs and configuration.
- Run a game day: kill a RegionServer under load and time the WAL split and reassignment.