HBase is built to absorb writes: a Put is appended to a log, inserted into a sorted in-memory buffer, and acknowledged, with all the expensive work of writing sorted files deferred to the background. When writes are slow, the cause is almost never the foreground path being inherently expensive. It is one of two things: a single resource in the foreground path, usually the write-ahead log sync, is saturated; or background work has fallen behind and the RegionServer is deliberately blocking writers until it catches up.
This page is a diagnostic procedure, layer by layer, in the same spirit as HBase read performance. The architecture of the write path itself is covered in the WAL and MemStore write path; here the question is how to find which layer is limiting you and what to change. Numeric defaults quoted below come from the shipped hbase-default.xml for the 2.x line; confirm them against your version.
First, say which write problem you have
Three symptoms look alike on a dashboard and have different causes, so name yours before tuning anything.
- Low throughput with flat latency. The cluster is idle and the client is not sending enough: too few threads, tiny batches, or every write going to one region.
- High p99 put latency all the time. A foreground resource is saturated, most often WAL sync to HDFS, sometimes the RPC handler pool or garbage collection.
- Periodic stalls. Latency is fine, then writes to some regions stop for seconds to minutes. That is backpressure: MemStores hit their blocking size or stores hit the file-count limit while flushes and compactions catch up.
Layer 1: the client decides your ceiling
A single synchronous table.put(put) per row pays a full network round trip and a WAL sync per row. Throughput then equals threads divided by latency: 8 threads at 5 ms is 1,600 rows per second no matter how big the cluster is. Batching amortises both costs, because the RegionServer appends the whole batch to the WAL and syncs once.
For ingest, use BufferedMutator, which accumulates mutations and sends them grouped by RegionServer when the buffer fills. The shipped hbase.client.write.buffer is 2 MB (2097152 bytes); a few megabytes to tens of megabytes suits bulk streams, while latency-sensitive writers should keep it small or flush explicitly. Errors from buffered writes arrive asynchronously, so always install an exception listener, or failures vanish.
// Throughput-oriented ingest: one shared Connection, a BufferedMutator per writer thread pool
BufferedMutatorParams params = new BufferedMutatorParams(TableName.valueOf("events"))
.writeBufferSize(8L * 1024 * 1024) // flush to servers every ~8 MB
.listener((e, mutator) -> { // async failures surface here, not at mutate()
for (int i = 0; i < e.getNumExceptions(); i++) {
log.error("put failed row={} server={}", Bytes.toStringBinary(e.getRow(i).getRow()),
e.getHostnamePort(i), e.getCause(i));
deadLetter.add(e.getRow(i));
}
});
try (BufferedMutator mutator = connection.getBufferedMutator(params)) {
for (Event ev : source) {
Put put = new Put(rowKey(ev)); // salted or hashed prefix, never a raw timestamp
put.addColumn(CF, Q_PAYLOAD, ev.payload());
mutator.mutate(put); // buffered; returns immediately
}
mutator.flush(); // push the tail before reporting success
}The second client decision is the row key. A key that starts with a timestamp or a sequence number sends every write to the last region, so one RegionServer's WAL and one MemStore take the whole load while the rest of the cluster idles. Prefix the key with a small hash bucket and pre-split the table on those buckets, accepting that time-range scans must fan out across buckets. Key design is covered fully in HBase hotspotting.
// Spread a monotonically increasing key across 16 pre-split regions
static byte[] rowKey(Event ev) {
int bucket = (ev.deviceId().hashCode() & 0x7fffffff) % 16;
return Bytes.add(new byte[] { (byte) bucket },
Bytes.toBytes(ev.deviceId()),
Bytes.toBytes(Long.MAX_VALUE - ev.timestampMillis()));
}For the asynchronous Java client, the same principles hold: issue many outstanding requests and batch them, and bound concurrency so a slow server produces backpressure rather than an unbounded queue in your process.
Layer 2: RPC handlers and the call queue
Each RegionServer serves requests from a pool of handler threads, 30 by default (hbase.regionserver.handler.count). A handler holds its thread for the whole Put, including the WAL sync, so handlers are consumed by waiting on HDFS, not by CPU. If every handler is busy, new requests queue, and the RegionServer's queue-time metrics rise while processing time stays flat.
Read the split directly: the server exposes queue call time and processing call time histograms along with the number of calls waiting in the general queue. High queue time with modest processing time means the pool is too small for the concurrency, and raising the handler count helps. High processing time means the handlers are waiting on something below them, and adding handlers only adds more waiters; go to layer 3.
Layer 3: the WAL sync is usually the bottleneck
Every durable write is appended to the RegionServer's write-ahead log and made durable through the HDFS pipeline before it is acknowledged. The default provider in 2.x writes with an asynchronous output stream and groups concurrent appends into one sync, so throughput depends on how many writes each sync carries and how long each sync takes. A sync crosses the network to three DataNodes and, depending on configuration, their disks, so its latency is set by the slowest DataNode in the pipeline.
Symptoms: WAL sync time percentiles rise, slow sync counts increase in the RegionServer log, and put latency tracks them. Fixes, in order of preference:
- Fix the slow DataNode. A single degraded disk or saturated network link on one DataNode inflates every pipeline that includes it; WAL sync latency is a sensitive detector of bad hardware.
- Increase batching at the client so each sync carries more data.
- Use more than one WAL per RegionServer with
hbase.wal.providerset tomultiwal, which spreads regions across several logs and therefore several pipelines, trading more files and slower recovery for parallel syncs. - Put the WAL on faster storage using HDFS storage policies, if your platform supports it.
- Relax durability per table or per mutation only where losing the most recent writes on a crash is acceptable. The levels and exactly what each risks are covered in WAL durability levels; skipping the WAL also skips replication for those edits.
Layer 4: MemStore limits and why regions block
After the WAL, a write is inserted into the MemStore of its column family, which is an in-memory sorted structure. That insert is cheap. The limits around it are what hurt. Each region flushes when its MemStores reach hbase.hregion.memstore.flush.size, 128 MB by default. If writes arrive faster than flushes complete, the region blocks new updates at flush size times hbase.hregion.memstore.block.multiplier, 4 by default, so 512 MB, and clients receive RegionTooBusyException and retry.
There is also a server-wide limit. The total of all MemStores on a RegionServer is capped at a fraction of heap set by hbase.regionserver.global.memstore.size (the shipped description gives 0.4 of heap when unset), and when usage reaches the lower limit the server starts forcing flushes of the largest regions; if it reaches the upper limit, all writes on the server block until flushing brings it down. Server-wide blocking is the most dramatic stall you will see: every region on the host stops at once.
Column families multiply all of this. Each family in each region has its own MemStore and its own set of files, so a write-heavy family next to quiet ones still adds per-store overhead, and which families a region flush includes depends on the configured flush policy. Check that policy for your version, and keep write-heavy data in as few families as possible.
Layer 5: flush and compaction backpressure
Each flush creates a new HFile in each store. Compaction merges them in the background. If a store accumulates hbase.hstore.blockingStoreFiles files, 16 by default, further flushes of that region are delayed until compaction reduces the count or hbase.hstore.blockingWaitTime (90 seconds by default) expires. While flushes are delayed the MemStore keeps growing toward its blocking size, so a compaction backlog becomes a write stall one step later.
The chain is worth memorising, because the symptom appears far from the cause: compaction falls behind, the file count reaches 16, flushes wait, MemStores reach 512 MB, updates block, and clients see latency spikes and RegionTooBusyException. The metrics that expose it are the flush queue length, the compaction queue length, store file counts per region, and the time updates were blocked. Only 2 flusher threads run by default (hbase.hstore.flusher.count); on hosts with many actively written regions and fast disks, more flushers help if the flush queue grows while disks are idle.
<!-- hbase-site.xml: values shown are the shipped defaults unless marked -->
<property><name>hbase.hregion.memstore.flush.size</name><value>134217728</value></property> <!-- 128 MB -->
<property><name>hbase.hregion.memstore.block.multiplier</name><value>4</value></property>
<property><name>hbase.hstore.blockingStoreFiles</name><value>16</value></property>
<property><name>hbase.hstore.blockingWaitTime</name><value>90000</value></property>
<property><name>hbase.hstore.flusher.count</name><value>2</value></property>
<property><name>hbase.regionserver.handler.count</name><value>30</value></property>
<!-- changed for a write-heavy RegionServer after measurement: -->
<property><name>hbase.wal.provider</name><value>multiwal</value></property>Raising blockingStoreFiles hides the stall by allowing more files, at the cost of slower reads and bigger compactions later. It is a legitimate lever for short bursts, not a fix for sustained ingest that compaction cannot keep up with. Compaction policy and throughput are covered in HBase compaction.
Worked example: sizing a RegionServer for ingest
A RegionServer has a 32 GB heap and hosts 200 regions of a table with one column family, all receiving writes because the keys are hashed. The global MemStore limit is 0.4 x 32 = 12.8 GB. Spread across 200 active regions that is 64 MB each, half of the 128 MB flush size, so regions never reach their own flush threshold. Instead the server repeatedly hits its global lower limit and force-flushes the largest regions at around 64 MB.
At 100 MB per second of ingest to this server, it produces a new file roughly every 0.64 seconds across its regions, about 1.6 files per second, twice as many files as the same data flushed at 128 MB. Compaction must rewrite that data several times, the store file counts climb, and after a few hours of sustained load regions start hitting the 16-file limit and stalling.
The fixes follow from the arithmetic. Reduce actively written regions per server by using fewer, larger regions (the default maximum region size is 10 GB) or by adding servers; 80 active regions would each get 160 MB of the global budget and flush at full size. Give the MemStore more heap if reads can afford a smaller block cache. And if the load is really a bulk import rather than a live stream, stop writing through the RegionServers at all.
When not to use the write path: bulk load
Batch ingest of historical data through Puts pays the WAL, MemStore, flush and compaction costs for data that is already complete and sorted. Bulk load writes HFiles directly with a MapReduce or Spark job and hands them to RegionServers, which adopt them atomically without touching the WAL or MemStore. Throughput is limited by the job, not the cluster's write path. The trade-offs are that bulk-loaded data is not in the WAL, so replication needs its own handling, and that files must match region boundaries. See HBase bulk load.
Failure modes and first responses
| Symptom | Likely cause | First response |
|---|---|---|
| All writes to one server slow, others idle | Hot region from sequential keys | Salt or hash keys; pre-split |
| p99 latency tracks WAL sync time | Slow DataNode or small batches | Find the slow node; batch more; consider multiwal |
| High queue time, low processing time | Handler pool too small | Raise handler count |
RegionTooBusyException on some regions | MemStore at blocking size | Check flush queue and store file counts |
| Every region on a host stops at once | Global MemStore limit | Fewer active regions per server or more MemStore heap |
| Stalls after hours of steady load | Compaction backlog reaching blockingStoreFiles | Increase compaction throughput or reduce file creation |
| Long pauses with no queue growth | Garbage collection | Check GC logs; MSLAB is on by default for this reason |
What to do next
- Classify the problem as throughput, steady latency or periodic stalls before changing any setting.
- Check the client: batching with BufferedMutator, an exception listener, enough concurrency, and a row key that spreads load.
- On the RegionServer, compare queue time with processing time, then WAL sync time with put latency.
- Graph flush queue, compaction queue, store files per store and blocked update time together; stalls show up there first.
- Compute global MemStore budget divided by actively written regions and compare it with the 128 MB flush size.
- Move historical or batch imports to bulk load, and relax durability only for data you can afford to lose.