An HBase cluster tells you most of what is wrong with it in its own log files, but it tells you in pieces. A slow write shows up as a WAL sync warning on one RegionServer, a long garbage-collection pause on the same host, a burst of slow-RPC entries a few seconds later, and client retries somewhere else entirely. Reading those files one host at a time with ssh and grep works for a five-node cluster and fails at fifty, because the correlation across hosts and time is the whole point.
This article is about aggregating HBase's own daemon logs, the HMaster and RegionServer logs, for operations. It is not about using HBase as a log store, and it is not about YARN container log aggregation. It covers what HBase writes and where, the handful of log lines worth turning into signals, a shipping and parsing pipeline, a truncation trap in the default layout that silently breaks JSON parsing, a worked incident investigation, the online slow log as a complement, and retention and level settings that keep volume sane without throwing away the evidence. Configuration names and defaults are taken from the HBase 2.6 source.
What HBase writes, and where
The start scripts decide file names and locations. bin/hbase-daemon.sh writes to $HBASE_LOG_DIR, which defaults to $HBASE_HOME/logs, using the prefix hbase-$HBASE_IDENT_STRING-$command-$HOSTNAME, where the ident string defaults to the user running the daemon and the command is master, regionserver and so on. Each daemon produces three files with that prefix:
.log: the application log written by the logging framework. This is the one to aggregate..out: the process's standard output and error, which catches JVM crash messages and anything printed before logging initialises..gc: the JVM garbage-collection log, if GC logging is enabled inhbase-env.sh.
HBase 2.5 moved from log4j 1 to log4j2, so the configuration file is conf/log4j2.properties on 2.5 and later and conf/log4j.properties on 2.4 and earlier. Many fleets run both during an upgrade, and a pipeline built for one breaks on the other. The daemon script sets the root logger from HBASE_ROOT_LOGGER, which defaults to INFO,RFA: INFO level to the size-based rolling file appender. The relevant part of the stock 2.6 configuration is:
appender.RFA.type = RollingFile
appender.RFA.fileName = ${sys:hbase.log.dir:-.}/${sys:hbase.log.file:-hbase.log}
appender.RFA.filePattern = ${sys:hbase.log.dir:-.}/${sys:hbase.log.file:-hbase.log}.%i
appender.RFA.layout.pattern = %d{ISO8601} %-5p [%t] %c{2}: %.1000m%n
appender.RFA.policies.size.size = ${sys:hbase.log.maxfilesize:-256MB}
appender.RFA.strategy.max = ${sys:hbase.log.maxbackupindex:-20}Two consequences matter for shipping. Rotation renames the live file to an indexed name such as hbase.log.1 and reuses the base name, so your agent must follow files by inode or identity, not by name, or it will either miss the tail of a rotated file or re-read it. And the defaults keep up to 21 files of 256 MB per daemon, about 5 GB on local disk; if the shipper falls behind by longer than that window, the oldest unshipped data is deleted before it leaves the host. There is also a security audit appender, writing SecurityAuth.audit, which belongs in a separate, access-controlled stream.
The log lines that carry the signal
Most of an INFO log is routine: region opens, flushes, compactions. A small set of lines carries most of the operational signal, and each has a configurable threshold:
| Log line starts with | Level | Emitted when | Threshold key and default |
|---|---|---|---|
(responseTooSlow): {...} | WARN | An RPC took longer than the threshold | hbase.ipc.warn.response.time, 10000 ms |
(responseTooLarge): {...} | WARN | An RPC response exceeded the size threshold | hbase.ipc.warn.response.size, 100 MiB |
Slow sync cost: N ms, current pipeline: [...] | INFO | One WAL sync exceeded the threshold | hbase.regionserver.wal.slowsync.ms, 100 ms |
Detected pause in JVM or host machine (eg GC) | INFO, then WARN | The JVM or host stalled | jvm.pause.info-threshold.ms 1000, jvm.pause.warn-threshold.ms 10000 |
Region is too busy due to exceeding memstore size limit. | WARN | Writes blocked; carries a RegionTooBusyException with Over memstore limit= | memstore blocking size |
The responseTooSlow and responseTooLarge entries are the richest: after the tag, the RegionServer writes a JSON object with fields including processingtimems, queuetimems, responsesize, client, method and the call parameters. Queue time separates an overloaded handler pool from a slow operation. The WAL sync line names the HDFS DataNodes in the write pipeline, which is how you find one bad disk among hundreds. The WAL also rolls itself when a single sync exceeds hbase.regionserver.wal.roll.on.sync.ms (10000 ms) or when 100 slow syncs, hbase.regionserver.wal.slowsync.roll.threshold, accumulate within a minute, so a run of slow-sync lines followed by a roll is a known pattern, not a coincidence.
Note the levels. Two of the most useful lines are INFO. Setting the root logger to WARN to cut volume, a common first reaction to log cost, removes slow-sync evidence and short JVM pauses entirely.
Pipeline architecture
The architecture is the standard one for any JVM fleet. An agent on every node tails the files, remembers its offset per file across restarts, and ships batches over the network. Fluent Bit, Vector, Filebeat and the OpenTelemetry Collector's file receiver all do this; pick the one your organisation already operates. A parsing stage, in the agent or centrally, turns each event into fields: timestamp, level, thread, logger, message, plus host, cluster and daemon role taken from the file name. The store keeps raw events for search and feeds derived counters to your metrics system, because alerts on log-derived rates are cheaper and more stable than alerts on full-text queries.
The one HBase-specific requirement is multiline handling. Exceptions span dozens of lines, and only the first carries a timestamp. The rule is simple: a new event starts at a line that begins with a timestamp, and every other line belongs to the previous event. Configure that in the agent, not downstream, because once lines are split into separate events their order across a busy pipeline is not guaranteed.
Parsing the default layout
A parser for the default pattern needs to handle the ISO8601 timestamp (log4j2 writes a T between date and time and a comma before milliseconds), a padded level, a bracketed thread name, an abbreviated logger, and the message. The following Python is the same logic most agents express in their own configuration language; it is also useful on its own for incident triage over a directory of copied logs:
import json, re, sys
HEAD = re.compile(
r"^(?P<ts>\d{4}-\d{2}-\d{2}[T ]\d{2}:\d{2}:\d{2},\d{3}) "
r"(?P<level>[A-Z]+)\s+\[(?P<thread>[^\]]*)\] "
r"(?P<logger>[^:]+): (?P<msg>.*)$")
SLOW_RPC = re.compile(r"^\((?P<tag>response[A-Za-z &]+)\): (?P<body>\{.*)$")
SLOW_SYNC = re.compile(r"^Slow sync cost: (?P<ms>\d+) ms, current pipeline: (?P<pipe>.*)$")
def events(lines):
cur = None
for line in lines:
m = HEAD.match(line.rstrip("\n"))
if m:
if cur:
yield cur
cur = m.groupdict()
cur["extra"] = []
elif cur:
cur["extra"].append(line.rstrip("\n")) # stack trace lines
if cur:
yield cur
def classify(ev):
msg = ev["msg"]
if len(msg) == 1000:
ev["maybe_truncated"] = True # %.1000m cap reached
if (m := SLOW_RPC.match(msg)):
try:
ev["rpc"] = json.loads(m["body"])
except ValueError:
ev["rpc_parse_error"] = True # see truncation trap
ev["kind"] = m["tag"]
elif (m := SLOW_SYNC.match(msg)):
ev["kind"], ev["sync_ms"], ev["pipeline"] = "slow_sync", int(m["ms"]), m["pipe"]
elif msg.startswith("Detected pause in JVM"):
ev["kind"] = "jvm_pause"
return ev
for ev in map(classify, events(open(sys.argv[1], encoding="utf-8"))):
if "kind" in ev:
print(ev["ts"], ev["kind"], ev.get("sync_ms") or ev.get("rpc", {}).get("processingtimems"))
The truncation trap in the default pattern
Look again at the stock pattern: %.1000m. In log4j2's pattern layout, a precision after the dot is a maximum width, and when a value is longer the extra characters are removed from the beginning, not the end. A responseTooSlow entry for a large multi-get or scan, with call parameters serialised into the JSON, easily exceeds 1,000 characters. When it does, the log keeps the last 1,000: the (responseTooSlow): { prefix and the leading fields are gone, the line no longer matches any slow-RPC rule, and the event is silently classified as an unremarkable INFO-looking fragment. The worst RPCs are exactly the ones you lose.
There are three fixes. Change the precision to %.-1000m, where the minus after the dot truncates from the end, which keeps the tag and the timing fields and drops the tail of the parameters. Raise the limit, accepting larger files. Or stop relying on the log line for full detail and enable the online slow log, described below, which keeps complete records. Whichever you choose, count the two flags the parser above sets: messages exactly 1,000 characters long, which hit the cap in either direction, and slow-RPC entries whose JSON fails to parse. A sudden rise in either is how you learn that lines are being cut or that someone changed the layout.
Worked example: finding a slow DataNode
Here is how aggregated logs shorten a real investigation. At 14:05 the application team reports p99 write latency jumping from 20 ms to 900 ms for ten minutes. Metrics show the cluster-wide write rate was normal and no region was blocked. With logs in one place, the triage runs as three queries over 13:55 to 14:20:
- Count slow-sync events per RegionServer per minute. Seven of forty RegionServers show bursts, all starting at 14:03.
- Extract the pipeline field from those events and count DataNode appearances. One DataNode appears in every slow pipeline across all seven RegionServers.
- Check JVM-pause and responseTooSlow events on the same seven hosts. There are no long pauses, so GC is ruled out, and the slow RPCs have high processing time but low queue time, which points at the write path rather than an exhausted handler pool.
The conclusion, one slow DataNode disk dragging every WAL pipeline that included it, takes a few minutes with aggregation and much longer host by host, because no single RegionServer's log looks alarming on its own. The fix is on the HDFS side; the follow-up for HBase is an alert on the per-minute slow-sync rate per host and a dashboard of the top pipeline members, both derived from the same parsed field.
The online slow log
HBase also keeps its own record of slow and large RPCs, separate from the log files. With hbase.regionserver.slowlog.buffer.enabled set to true (the default is false), each RegionServer keeps an in-memory ring buffer of the most recent slow and large calls, sized by hbase.regionserver.slowlog.ringbuffer.size (default 256). Unlike the log line, the in-memory record is complete. Setting hbase.regionserver.slowlog.systable.enabled as well persists entries to the hbase:slowlog system table. The shell reads the buffers:
hbase> get_slowlog_responses '*', {'LIMIT' => 50}
hbase> get_slowlog_responses '*', {'TABLE_NAME' => 'orders', 'CLIENT_IP' => '10.0.4.17', 'FILTER_BY_OP' => 'AND'}
hbase> get_largelog_responses '*', {'LIMIT' => 20}
hbase> clear_slowlog_responsesTwo shell details catch people out: the default limit is 10 records per server, and multiple filters are combined with OR unless you pass 'FILTER_BY_OP' => 'AND', so a filter by table and client returns far more than expected. Treat the slow log as the detailed drill-down and the aggregated log stream as the cross-host timeline; you want both.
Volume, levels and retention
Keeping the pipeline useful is mostly about volume and levels:
- Lower noise per logger, not at the root. Leave the root at INFO and raise specific chatty loggers to WARN in
log4j2.properties, keepingorg.apache.hadoop.hbase.regionserver.walandorg.apache.hadoop.hbase.util.JvmPauseMonitorat INFO. - Size local retention to shipper outages. Set
-Dhbase.log.maxfilesizeand-Dhbase.log.maxbackupindexthroughHBASE_OPTSso local files cover the longest pipeline outage you expect, and alert on shipper lag. - Tier central retention. Keep raw events searchable for days to weeks, and keep derived per-minute counts for months; capacity planning and regression hunts need the long history only in aggregate.
- Tag every event. Add cluster, daemon role and HBase version at the agent. Mixed-version fleets during upgrades need the version tag to pick the right parser.
- Do not log DEBUG cluster-wide. Enable DEBUG on one logger, on one host, for a bounded time, and revert it through configuration management.
Failure modes
The failure modes of log aggregation itself are worth listing, because each hides an incident:
- Front-truncated JSON from the default
%.1000mpattern drops the slowest RPCs from your counts. - Rotation races in agents that follow files by name lose or duplicate the tail of each rotated file.
- Multiline splitting scatters stack traces into separate events, so the exception class never sits next to the message that explains it.
- Clock skew across hosts reorders the cross-host timeline; keep NTP healthy and prefer the event timestamp over ingest time.
- Backpressure on the node. An agent with unbounded buffers can consume memory on a RegionServer host during a store outage; cap buffers and spill to disk.
- Level changes that silence INFO remove the slow-sync and short-pause evidence without any error.
Related reading
Related reading: HBase alerting covers which conditions should page, HBase metrics in depth covers the JMX metrics that complement log-derived counts, structured logging practices covers field design, Loki versus Elastic versus ClickHouse compares log stores, and Hadoop log aggregation covers the YARN side of the same cluster.
What to do next
- Inventory HBase versions in the fleet and note which hosts use log4j.properties and which use log4j2.properties.
- Deploy a node agent that tails the .log, .out and .gc files by file identity, with multiline rules keyed on the leading timestamp.
- Change the layout precision from %.1000m to %.-1000m, or raise it, and add a counter for unparseable slow-RPC JSON.
- Extract responseTooSlow, responseTooLarge, Slow sync cost and JVM pause events into fields, and publish per-host per-minute counts.
- Enable the online slow log ring buffer on RegionServers, and decide whether to persist it to hbase:slowlog.
- Keep the root logger at INFO, quiet specific loggers instead, and size local rotation to cover shipper outages.
- Rehearse one investigation, such as finding a slow DataNode from WAL pipelines, before you need it.