Every write to HBase lands in a MemStore, an in-memory sorted buffer per column family per region, and stays there until a flush writes it out as an HFile. MemStore tuning is the business of deciding how much memory those buffers get, when they flush, and what happens when writes outrun flushing. Get it right and flushes produce large files that compaction handles cheaply. Get it wrong and the server writes thousands of tiny files, compaction falls behind, and writes stall even though the disks and CPUs look idle.
What a flush does internally, and why the MemStore is a skip list, is covered in MemStore and flushes. This page is about the settings. It goes through each budget that can trigger a flush, gives the HBase 2.6 default for each, shows how they interact, and works through the most common real failure: too many active regions sharing one global budget.
Four budgets, one flush
A MemStore flushes when the tightest of several independent budgets runs out. The global budget caps the total MemStore memory on a RegionServer. The per-region budget caps one region. The per-family policy decides which stores in a region are written when the region flushes. And two housekeeping triggers, a timer and the count of write-ahead log files, flush regions that are neither big nor under pressure. The fix depends on which one is firing.
<!-- hbase-site.xml: the MemStore knobs, shown at their HBase 2.6 defaults -->
<property><name>hbase.regionserver.global.memstore.size</name><value>0.4</value></property>
<property><name>hbase.regionserver.global.memstore.size.lower.limit</name><value>0.95</value></property>
<property><name>hfile.block.cache.size</name><value>0.4</value></property>
<property><name>hbase.hregion.memstore.flush.size</name><value>134217728</value></property>
<property><name>hbase.hregion.memstore.block.multiplier</name><value>4</value></property>
<property><name>hbase.hregion.percolumnfamilyflush.size.lower.bound.min</name><value>16777216</value></property>
<property><name>hbase.regionserver.optionalcacheflushinterval</name><value>3600000</value></property>
<property><name>hbase.hstore.flusher.count</name><value>2</value></property>
<property><name>hbase.hstore.blockingStoreFiles</name><value>16</value></property>
<property><name>hbase.hregion.memstore.mslab.enabled</name><value>true</value></property>
<property><name>hbase.hregion.memstore.mslab.chunksize</name><value>2097152</value></property>
<property><name>hbase.hregion.compacting.memstore.type</name><value>NONE</value></property>These are the values in HBase 2.6's hbase-default.xml and source, for the two global fractions documented there as 0.4 and 0.95. Change them one at a time, watching the metrics described below.
The global watermarks
hbase.regionserver.global.memstore.size is the fraction of the RegionServer heap that all MemStores together may use, 40 percent by default. It is the upper watermark: when the total reaches it, the server blocks every update until flushes bring it down. That is a server-wide write stall, so the second setting exists to avoid ever reaching it. hbase.regionserver.global.memstore.size.lower.limit is a fraction of the upper watermark, 95 percent by default; above it, the server starts forcing flushes, picking the largest regions first, while still accepting writes.
This budget competes with the block cache. hfile.block.cache.size also defaults to 40 percent of heap, and HBase refuses to start if the two together leave less than 20 percent of the heap for everything else: RPC buffers, scanners, compaction and the JVM itself. Raising the MemStore fraction therefore means lowering the on-heap block cache, or moving the cache off-heap. For a write-heavy table whose reads come mostly from recent data, 0.5 MemStore and 0.3 cache is a common shift; for a read-heavy table, leave the default.
Per region: flush size and the block multiplier
hbase.hregion.memstore.flush.size is the region-level trigger: when the MemStores of one region together exceed 128 MB, that region is queued to flush. Flushes run on hbase.hstore.flusher.count threads, two by default, so under heavy load the queue grows and regions keep accepting writes past 128 MB while they wait.
The block multiplier is the region's safety valve. When a region's MemStore reaches hbase.hregion.memstore.block.multiplier times the flush size, 4 x 128 MB = 512 MB by default, the region rejects writes with RegionTooBusyException until its flush completes. Clients retry, so users see latency before errors. It means the region ingests faster than one flush drains, or its flush is queued behind others. Raising the multiplier only hides this; add flusher threads or fix what slows flushes.
Do not raise the flush size by itself to get bigger files. If the global budget divided across your active regions is already below 128 MB, regions never reach the flush size: they are flushed earlier by global pressure, and the setting changes nothing.
The many-active-regions trap
The global budget is shared by every region currently receiving writes. With hundreds of active regions, each gets only a sliver of it, and pressure-driven flushes write small files long before any region reaches 128 MB. The calculator below makes the effect concrete for a 31 GB heap and 40 MB/s of ingest into one RegionServer. It approximates steady state by assuming active regions fill evenly, which is close for hashed or salted keys.
import math
MB = 1024 * 1024
def memstore_plan(heap_gb, global_frac, lower_frac, active_regions, cfs,
flush_mb=128, ingest_mb_s=40):
heap = heap_gb * 1024 * MB
upper = heap * global_frac # updates block above this
lower = upper * lower_frac # forced flushing starts here
per_region = lower / active_regions # what each region can hold
effective = min(per_region, flush_mb * MB)
per_store = effective / cfs # bytes per flushed file (pre-compression)
files_per_hour = ingest_mb_s * MB * 3600 / per_store
print(f"heap {heap_gb} GB, {active_regions} active regions x {cfs} CFs")
print(f" global upper {upper / MB:,.0f} MB, lower {lower / MB:,.0f} MB")
print(f" memstore per region before pressure flush: {per_region / MB:,.1f} MB")
print(f" typical flushed file: {per_store / MB:,.1f} MB "
f"({'size-triggered' if per_region >= flush_mb * MB else 'pressure-triggered'})")
print(f" new store files per hour across the server: {files_per_hour:,.0f}")
rounds = math.log(1024 * MB / per_store, 3) # 3-way merges up to a 1 GB file
print(f" per store per hour: {files_per_hour / (active_regions * cfs):,.1f} files; "
f"merge rounds to reach 1 GB: {rounds:.1f}")
memstore_plan(heap_gb=31, global_frac=0.4, lower_frac=0.95, active_regions=400, cfs=3)
memstore_plan(heap_gb=31, global_frac=0.4, lower_frac=0.95, active_regions=60, cfs=1)Running it prints:
heap 31 GB, 400 active regions x 3 CFs
global upper 12,698 MB, lower 12,063 MB
memstore per region before pressure flush: 30.2 MB
typical flushed file: 10.1 MB (pressure-triggered)
new store files per hour across the server: 14,325
per store per hour: 11.9 files; merge rounds to reach 1 GB: 4.2
heap 31 GB, 60 active regions x 1 CFs
global upper 12,698 MB, lower 12,063 MB
memstore per region before pressure flush: 201.0 MB
typical flushed file: 128.0 MB (size-triggered)
new store files per hour across the server: 1,125
per store per hour: 18.8 files; merge rounds to reach 1 GB: 1.9In the first configuration, 400 regions with three column families each, the server flushes 10 MB files, about 14,000 of them an hour. Each file must then be merged into larger ones by compaction; with merges of three files at a time, data is rewritten about four times on its way to a 1 GB file, against about twice in the second, which writes more files per store but far fewer in total. That extra rewriting is pure write amplification, and when compaction falls behind, a store passes hbase.hstore.blockingStoreFiles, 16 files by default, and the region stops accepting writes until compaction catches up.
Files on disk are smaller still, after compression. The fix is rarely one setting: reduce the regions active at once, through fewer, larger regions or keys that write to a moving window; remove column families that do not need separate storage; raise the global fraction if the read cache can afford it; or add RegionServers. How these choices fit into the full write path is covered in the write performance guide.
Column families and the flush policy
A region flush does not have to write every store. The default policy, set by hbase.regionserver.flush.policy, is FlushAllLargeStoresPolicy: when a region flushes, it writes only the stores above a lower bound, and falls back to flushing all of them when none qualifies. The bound is derived from the region's flush size divided by the number of families, with a floor of hbase.hregion.percolumnfamilyflush.size.lower.bound.min, 16 MB by default.
This stops a small family producing tiny files whenever its large sibling flushes. It does not make families independent: they still share the region's flush size, block multiplier and the global budget. If one family takes 95 percent of the writes, the others mostly ride along. Keep tables to one or two families unless access patterns genuinely differ.
Flushes nobody asked for: WAL count and the hourly timer
Two housekeeping triggers flush regions that are neither large nor under pressure. The first is the number of write-ahead log files. A WAL file can be archived only when every edit in it has been flushed, so if a rarely written region holds one old edit, it pins every WAL written since. When the count exceeds hbase.regionserver.maxlogs, the server forces flushes of the regions holding the oldest edits. In HBase 2.x the default is computed: twice the global MemStore size divided by the WAL roll size, with a minimum of 32, where the roll size is the WAL block size times a 0.5 multiplier. These forced flushes are typically small, and the log line reports too many WALs.
The second is hbase.regionserver.optionalcacheflushinterval, one hour by default: an edit that has sat in memory that long is flushed regardless of size. On tables with a long tail of lightly written regions, this produces hourly small flushes. Raising it lengthens recovery replay after a crash, so change it only alongside the WAL count. How WALs roll and are retained is covered in the WAL deep dive.
MSLAB: the memory cost of each active store
The MemStore-Local Allocation Buffer copies cell data into 2 MB chunks, hbase.hregion.memstore.mslab.chunksize, so that a flushed MemStore frees whole chunks instead of leaving millions of small holes for the garbage collector. Cells larger than hbase.hregion.memstore.mslab.max.allocation, 256 KB, bypass it. Chunks are recycled through a pool whose maximum, hbase.hregion.memstore.chunkpool.maxsize, defaults to 1.0, the whole global MemStore, with an initial size of 0.
The cost to know: every store with data holds at least one chunk, and the size counters track cells, not chunks. At 400 regions with three families, 1,200 stores pin at least 2.4 GB of heap the watermarks never see, so leave headroom. Leave MSLAB on; it is one of the main defences against long collector pauses, covered in the GC tuning guide. Reduce active stores instead.
In-memory compaction and off-heap MemStores
In-memory compaction, also called Accordion, applies the log-structured merge idea inside the MemStore. It is configured per family with the IN_MEMORY_COMPACTION attribute, or globally with hbase.hregion.compacting.memstore.type, which defaults to NONE in 2.6. BASIC flattens older in-memory segments into a more compact index without removing data, cutting overhead per cell. EAGER also merges segments and drops overwritten versions and expired cells while still in memory. ADAPTIVE is documented as experimental and chooses between them based on how many duplicates it sees.
# HBase shell: enable in-memory compaction for one family that sees heavy overwrites.
alter 'counters', {NAME => 'c', IN_MEMORY_COMPACTION => 'EAGER'}
# Check what the family now carries.
describe 'counters'The gain depends on duplication. A counter table that rewrites the same rows constantly can hold far more logical updates per flush with EAGER, so flushes are rarer and files smaller. An append-only event table has nothing to eliminate, and the extra CPU can lower throughput. Test with production-shaped writes before enabling it widely.
An off-heap MemStore moves the cell data out of the Java heap. hbase.regionserver.offheap.global.memstore.size sets its size in megabytes; the default of 0 means on-heap. The JVM's direct memory limit, set through HBASE_OFFHEAPSIZE in hbase-env.sh, must also be large enough to cover it plus any other off-heap users such as the bucket cache, or the server fails when it tries to allocate.
Reading the symptoms
Each budget leaves a fingerprint. A growing flush queue means flushing is too slow: check HDFS write latency and flusher threads. RegionTooBusyException in client logs points to a single hot region hitting its block multiplier. Server-wide stalls with log lines about blocking updates mean the global upper watermark was hit. Store file counts per store climbing toward 16 mean compaction is losing, usually because flushes are too small. And bursts of small flushes with a too-many-WALs log line mean WAL pressure.
Track per RegionServer: total MemStore size against the two watermarks, flush queue length, time spent with updates blocked, store files per store, and the distribution of flushed file sizes. Flushed file size is the best single metric: it shows which budget is in charge. Where these metrics live and how the server's memory tenants compete is covered in the RegionServer article.
Failure modes
- Raising flush size under global pressure: regions never reach it, so nothing changes.
- Raising the block multiplier to silence errors: the region holds more unflushed data, flushes take longer, and the global watermark is hit instead.
- Memstore plus block cache over 80 percent: the RegionServer refuses to start after a config push.
- Too many families: each family is a store with its own chunk and files, multiplying small flushes.
- Long-tail regions pinning WALs: forced small flushes and a growing WAL directory.
- Off-heap MemStore without direct memory: allocation failures under load.
What to do next
- Graph flushed file sizes per RegionServer for a week, and identify which budget produces most flushes.
- Count active regions and stores per server; if global MemStore divided by active regions is below the flush size, reduce active regions before touching any setting.
- Check that the MemStore and block cache fractions together leave at least 20 percent of heap, and shift them toward the workload you have.
- Alert on
RegionTooBusyException, blocked update time and store files per store approachinghbase.hstore.blockingStoreFiles. - Trial
EAGERin-memory compaction on one overwrite-heavy family and compare flush counts and CPU before rolling it out. - Change one knob at a time, and record the metric you expect it to move before you change it.