Most HBase tuning advice is written for one direction at a time: how to make reads fast, or how to make ingest fast. Production clusters rarely get to choose. A profile store takes a steady stream of updates while serving lookups; a time-series table ingests continuously while dashboards scan it; a nightly batch job scans everything while the online service carries on. The difficult incidents happen at the boundary, when a write burst makes reads slow or a scan makes writes block, and the single-direction procedures do not explain why.

This article is about that interaction. It names the four resources that reads and writes share inside a RegionServer, shows how each write-side mechanism degrades reads, gives the levers that isolate the two, and walks a worked example from symptom to fix. For the single-direction procedures, see HBase Read Performance and HBase Write Performance.

Four shared resources

One RegionServer, two workloads: where writes and reads collidePutsBufferedMutatorGets and Scanspoint and rangeRPC call queues + handlersshared unless read.ratio > 0JVM heapMemStoresglobal 0.4 of heapBlock cache (L1)0.4 of heapDisk and HDFSFlush + compactionwrite and rewrite filesHFile readscache misseswritesreadsflushCollision points1. heap: memstore and cache share a fixed budget2. files: more flushes mean more HFiles per read until compaction3. disks: compaction I/O competes with cache-miss reads4. handlers: slow scans or blocked puts hold threads the other side needs
Reads and writes meet in four places inside a RegionServer: the call queues, the heap split, the number of files a read must consult, and the disks.

Every RegionServer has a fixed heap, a set of RPC handler threads, a set of disks behind HDFS, and a JVM garbage collector. Writes land in MemStores and the write-ahead log, are flushed to new HFiles, and are later rewritten by compaction. Reads consult the MemStore, the block cache and, on a miss, the HFiles. So every write mechanism has a read-side cost, and the defaults are a compromise that suits neither a write-heavy nor a read-heavy server particularly well.

Write-side eventRead-side effect
MemStore growsless heap available for the block cache if you rebalance the split toward writes
Flush creates an HFileGets and Scans consult one more file until compaction merges it
Many small flushesmore files, more seeks per read, bloom filters loaded for each
Compaction rewrites filesdisk and network I/O compete with cache misses; cached blocks of old files become useless
Store reaches the blocking file countputs to that region stall, holding handlers that reads also need

The heap split and global MemStore pressure

Two fractions divide the heap. hbase.regionserver.global.memstore.size (default 0.4) caps all MemStores on the server, and hfile.block.cache.size (default 0.4) sizes the on-heap block cache. HBase refuses to start if the two together leave too little heap for everything else; the classic rule is that they must not exceed 0.8 of the heap. When total MemStore usage reaches the lower limit, 0.95 of the global cap by default, the server starts forcing flushes of the largest MemStores; if it reaches the cap, updates block until flushes catch up.

Shifting the split is the most direct read-write trade-off you have. A write-heavy server can move toward 0.5 MemStore and 0.3 cache; a read-heavy one the other way. A better answer for most mixed workloads is to move the bulk of the block cache off heap with BucketCache, so the heap split mostly decides MemStore room and small index and bloom blocks, while data blocks live in a large off-heap cache that writes cannot squeeze. Details are in HBase BucketCache.

The less obvious problem is the interaction between the global cap and the per-region flush size. Each MemStore flushes at hbase.hregion.memstore.flush.size (default 128 MiB) and blocks updates at that size times hbase.hregion.memstore.block.multiplier (default 4). If many regions take writes at once, the global cap is reached long before any region reaches its flush size:

GiB, MiB = 1024**3, 1024**2
heap = 32 * GiB
memstore_global = 0.4 * heap          # hbase.regionserver.global.memstore.size
flush_floor     = 0.95 * memstore_global   # ...global.memstore.size.lower.limit
block_cache     = 0.4 * heap          # hfile.block.cache.size
active_regions  = 200                 # regions taking writes, one column family each
flush_size      = 128 * MiB           # hbase.hregion.memstore.flush.size

wanted = active_regions * flush_size
print(f"global memstore {memstore_global/GiB:.1f} GiB, flushes forced above {flush_floor/GiB:.2f} GiB")
print(f"regions would like {wanted/GiB:.1f} GiB before size-triggered flushes")
print(f"average file under global pressure: {flush_floor/active_regions/MiB:.0f} MiB")
print(f"regions that fit at full flush size: {int(flush_floor // flush_size)}")

For a 32 GiB heap and 200 actively written regions, the script prints a global MemStore of 12.8 GiB with forced flushes above 12.16 GiB, while the regions would like 25.0 GiB before flushing on their own. Under sustained load the average flushed file is about 62 MiB, half the configured size, and only 97 regions fit at full flush size. Twice as many files means twice the compaction work and more files per read in the meantime. Fewer actively written regions per server, or a larger MemStore share, is often worth more than any read-side tuning. MemStore internals are covered in HBase MemStore Tuning.

Files, compaction and the read path

Between a flush and the compaction that merges its output, every read of the affected rows may consult one more file. Bloom filters let a Get skip files that cannot contain the row, so point reads tolerate file growth reasonably well; scans cannot use row blooms to skip files and must merge every file that overlaps the range. That is why a write burst often hurts scans and range-heavy dashboards first.

Compaction bounds the file count, at an I/O cost. With the default settings a minor compaction becomes eligible when a store has at least three files (the code default for hbase.hstore.compaction.min) and merges at most hbase.hstore.compaction.max (default 10). When a store reaches hbase.hstore.blockingStoreFiles (default 16), flushes for that region are held while compaction catches up, for up to hbase.hstore.blockingWaitTime (default 90,000 ms); meanwhile its MemStore cannot drain, and once it reaches the blocking size, puts block. Raising the blocking count lets writes continue during bursts but allows reads to see more files; lowering it protects reads by stalling writes. Neither is free.

Compaction throughput is limited by a pressure-aware controller whose bounds default to 50 MiB/s and 100 MiB/s per server (hbase.hstore.compaction.throughput.lower.bound and .higher.bound). It runs near the lower bound when stores have few files and moves toward the upper bound as they approach the blocking count. On disks that also serve cache misses, the upper bound is a read-latency setting as much as a write setting. Raise it when compaction queues grow and files per store climb; lower it when compaction visibly coincides with read p99 spikes and file counts stay comfortable. HBase Compaction Tuning covers policy choice.

Separating reads from writes in the call queues

By default all calls share call queues and handlers. A RegionServer runs hbase.regionserver.handler.count handlers (default 30); the number of queues is the handler count times hbase.ipc.server.callqueue.handler.factor (default 0.1). When puts block on a region at its file limit, or long scans hold handlers while they read from disk, the other workload waits behind them in the same queues.

Two settings separate them. hbase.ipc.server.callqueue.read.ratio (default 0) dedicates a fraction of the queues and their handlers to reads, the rest to writes. hbase.ipc.server.callqueue.scan.ratio (default 0) then splits the read queues into short reads (Gets) and long reads (Scans). A starting point for a mixed server:

<!-- hbase-site.xml: separate read and write queues, then split reads into gets and scans -->
<property><name>hbase.regionserver.handler.count</name><value>60</value></property>
<property><name>hbase.ipc.server.callqueue.handler.factor</name><value>0.1</value></property>
<property><name>hbase.ipc.server.callqueue.read.ratio</name><value>0.6</value></property>
<property><name>hbase.ipc.server.callqueue.scan.ratio</name><value>0.3</value></property>

This gives six queues, most of them for reads, with a minority of those reserved for scans. The point is not the exact numbers; it is that a flood of scans can only exhaust the scan handlers, and stalled puts can only exhaust write handlers, so Gets keep a path. The cost is lost pooling: an idle write handler cannot help a read backlog. Watch the per-type queue lengths after the change and move the ratios toward whichever side queues.

Client-side levers on both sides

The client controls how much each side asks of the server. On the write side, batch through BufferedMutator so each RPC carries many mutations and handlers spend less time per row; the default write buffer is 2 MiB. Durability is a lever too: ASYNC_WAL acknowledges before the WAL sync and removes sync latency from puts, at the cost of losing recently acknowledged writes if a server crashes; use it only for data you can replay. On the read side, the single most important setting for mixed clusters is setCacheBlocks(false) on full-table and batch scans, so a one-off pass over the whole table does not evict the hot set that online Gets depend on.

// Writes: batch on the client so each RPC carries many mutations.
BufferedMutatorParams params = new BufferedMutatorParams(TableName.valueOf("profile"))
    .writeBufferSize(4 * 1024 * 1024);
try (BufferedMutator mutator = connection.getBufferedMutator(params)) {
    for (Event e : events) {
        Put put = new Put(rowKey(e)).addColumn(CF, QUAL, e.payload());
        mutator.mutate(put);
    }
}   // close() flushes the remaining buffer

// Analytics scans: do not let a full-table pass evict the hot set.
Scan scan = new Scan()
    .addFamily(CF)
    .setCacheBlocks(false)          // blocks read by this scan are not cached
    .setCaching(500)                // rows per RPC; bounded by max result size
    .setMaxResultSize(4L * 1024 * 1024);
try (ResultScanner rs = table.getScanner(scan)) {
    for (Result r : rs) { process(r); }
}

Scanner caching defaults to effectively unlimited rows per RPC, bounded by hbase.client.scanner.max.result.size (default 2 MiB). Bigger batches mean fewer round trips but longer handler holds per call; for scan queues shared with interactive work, moderate batches keep latency even.

Measuring a mixed workload

Tune mixed workloads with a mixed benchmark. YCSB workload A (50 percent reads, 50 percent updates) and workload B (95 percent reads) are reasonable shapes; better still, replay a captured ratio of your own traffic. Run at a fixed target throughput rather than flat out, because the interesting behaviour is latency at production load, and run long enough for several flush and compaction cycles; ten-minute runs measure an empty MemStore and a fresh cache. Record per-operation p99 alongside server state, which the RegionServer exposes over its JMX endpoint:

import requests

def rs_snapshot(host, port=16030):
    beans = requests.get(f"http://{host}:{port}/jmx", timeout=5).json()["beans"]
    by = {b["name"]: b for b in beans}
    srv = by["Hadoop:service=HBase,name=RegionServer,sub=Server"]
    ipc = by["Hadoop:service=HBase,name=RegionServer,sub=IPC"]
    return {
        "memstore_bytes": srv.get("memStoreSize"),
        "flush_queue": srv.get("flushQueueLength"),
        "compaction_queue": srv.get("compactionQueueLength"),
        "store_files": srv.get("storeFileCount"),
        "regions": srv.get("regionCount"),
        "cache_hit_pct": srv.get("blockCacheExpressHitPercent"),
        "read_queue": ipc.get("numCallsInReadQueue"),
        "write_queue": ipc.get("numCallsInWriteQueue"),
        "scan_queue": ipc.get("numCallsInScanQueue"),
    }

# Files per region is the number to watch on mixed workloads:
# s = rs_snapshot("rs1.example.internal"); print(s["store_files"] / s["regions"])

Files per region is the number that ties the two workloads together: rising files per region with a growing compaction queue predicts read degradation minutes before p99 moves. A falling cache hit percentage during compaction or batch scans points at eviction. Growth in the write queue while the read queue stays flat says writes are blocking, usually on the store file limit.

Worked example: the 02:00 spike and the Monday stall

A profile service runs on 20 RegionServers with 32 GiB heaps. It takes 40,000 updates per second and serves 20,000 Gets per second with a p99 target of 20 ms. Every night at 02:00 a batch job scans the whole table and p99 Gets climb to 150 ms; on Mondays after a marketing send, puts occasionally stall for over a minute.

The nightly spike shows a block cache hit percentage falling from the high nineties to below 70 within minutes of the scan starting, while the read queue grows. The scan was reading through the cache and evicting the hot profiles. Setting setCacheBlocks(false) in the batch job and adding a scan ratio so its RPCs use their own handlers brought the 02:00 p99 back under 25 ms.

The Monday stalls show store files per region climbing from four to sixteen during the send, the compaction queue growing, and write-queue length spiking while reads are unaffected until the stall. Each server hosted about 300 actively written regions, so global MemStore pressure produced small flushes, as in the arithmetic above. Moving data blocks to an off-heap BucketCache let the MemStore share rise to 0.5 without shrinking the effective cache, which puts forced flushes above 15.2 GiB; merging regions to about 110 write-active regions per server (13.75 GiB at full flush size) brought every region under that budget. Raising the compaction throughput upper bound to keep pace with the send removed the remaining stalls.

Failure modes

  • Tuning one direction under a one-direction benchmark. A write-only load test will happily recommend a split that ruins reads. Benchmark the mix.
  • Raising blockingStoreFiles to hide stalls. Writes stop stalling and reads quietly slow down as files accumulate. Fix flush size and compaction throughput first.
  • Too many actively written regions. Global pressure turns every flush into a small file. Count write-active regions per server, not total regions.
  • Batch scans through the cache. One job can evict the hot set for the whole server. Disable block caching for them.
  • Read and write queues split without measurement. Fixed ratios that do not match the traffic leave handlers idle on one side while the other queues.
  • Large heaps without GC review. Growing the heap to fit both MemStore and cache can lengthen pauses that hit both workloads at once; prefer off-heap cache.

What to do next

  1. Benchmark your real read-write mix at production throughput for at least an hour, recording per-operation p99.
  2. Export files per region, compaction queue, cache hit percentage and per-type queue lengths to your dashboards.
  3. Count write-active regions per server and compare their total flush size with the global MemStore cap.
  4. Move data blocks off heap with BucketCache before rebalancing the heap split.
  5. Set setCacheBlocks(false) on every batch and full-table scan.
  6. Enable read and scan queue ratios, then adjust them from measured queue lengths.
  7. Tune compaction throughput bounds against read latency, not only compaction backlog.
  8. Re-run the mixed benchmark after each change, one change at a time.
Key takeaway: Reads and writes in HBase share the heap, the files a read must consult, the disks and the RPC handlers, so tuning one side changes the other. Keep write-active regions within the global MemStore budget, move data blocks off heap, bound files per store with compaction throughput rather than a higher blocking limit, isolate scans and writes in their own call queues, keep batch scans out of the cache, and measure everything under the real mix.