The BlockCache is the RegionServer's read cache. It holds HFile blocks, the units HBase reads from disk: data blocks with cells, index blocks that say which data block holds a row, and bloom filter blocks that say whether a file can contain a row at all. A Get served from the cache costs microseconds. A miss costs a disk or network read, often several, and that is where most tail latency on read-heavy HBase clusters comes from.

This article is a tuning guide. How the caches are built is covered in HBase block cache architecture and HBase BucketCache. Here you get the memory budget and the startup check that enforces it, a method for sizing L1 and L2, the eviction and bucket-size details that decide whether the memory is used well, the per-table and per-request controls, the metrics that tell you whether a change worked, and a worked example. Names and defaults were checked against the HBase branch-2 source on 2026-10-04.

What the cache holds and why L1 and L2 differ

Start with the questions the cache answers on every read. For each HFile a Get must consult, HBase needs that file's bloom chunk (to skip the file if the row cannot be in it), its index blocks (to find the data block) and finally the data block. Index and bloom blocks are touched by every read that reaches a file, so they are small and extremely hot. Data blocks are large in total and their hotness follows the workload.

That split explains the standard layout. With only an on-heap LruBlockCache, all three types share one LRU. Once you configure a BucketCache, HBase 2 always runs a CombinedBlockCache: index and bloom blocks stay in the on-heap LRU, called L1, and data blocks go to the BucketCache, called L2. The source logs a warning if you try to turn combined mode off, because since 2.0 it is the only mode. The practical rule is that L1 must hold every index and bloom block of every open file, and L2 should hold the hot data set.

The memory budget and the startup check

Two fractions of the heap are set in hbase-site.xml: hbase.regionserver.global.memstore.size (default 0.4) for MemStores and hfile.block.cache.size (default 0.4) for L1. For many releases the rule has been that their sum must not exceed 0.8, and a RegionServer that violates it refuses to start. Recent branch-2 code expresses the same rule as a free-heap reserve, hbase.regionserver.free.heap.min.memory.size, defaulting to 20% of heap, and adds hfile.block.cache.memory.size to set L1 as an absolute size. Check which form your version has; the arithmetic is the same.

Where a RegionServer's memory goes, and which knob sizes each partMachine RAM, 128 GB in the worked exampleJava heap (-Xmx 31g)MemStoreglobal.memstore.size 0.4L1 LruBlockCachehfile.block.cache.sizeEverything elsekeep at least 20% freeDirect memory (off-heap)L2 BucketCachebucketcache.size (MB)RPC and read buffersOS page cache, DataNodeand the kernelL1: index and bloomsmall, hot, must not missL2: data blockslarge, sized to the hot setNot cachedfull scans, BLOCKCACHE falseWith a BucketCache configured, HBase 2 runs a CombinedBlockCache: index and bloom blocks in L1, data blocks in L2.Startup fails if memstore plus L1 leaves less than the free-heap reserve (20% of heap by default).
Heap holds MemStore and L1; direct memory holds L2. The page cache gets what is left, and it still matters for misses.

L2 lives outside the heap and is sized by hbase.bucketcache.size together with hbase.bucketcache.ioengine. Current branch-2 reads the size as megabytes and rejects anything under 1; older releases also accepted a value below 1 as a fraction, so do not copy settings between versions without checking. The engine is offheap for direct memory, or file:/path, files:, mmap: or pmem: for storage-backed caches. An off-heap L2 must fit inside the JVM's direct memory limit together with RPC buffers; that budget, HBASE_OFFHEAPSIZE, is worked through in HBase Off-Heap Memory, in depth.

Why not a huge heap and a huge L1? Every cached block on the heap is an object the collector must trace and, under churn, copy; off-heap L2 takes data blocks out of its view.

Sizing L1 and L2

Size the two tiers separately, in this order.

  1. L1 from the index and bloom footprint. The RegionServer metrics staticIndexSize and staticBloomSize report the index and bloom bytes of open store files. Set L1 to at least their sum plus 50% for growth, region moves and the leaf-index blocks loaded on demand. If L1 is too small, index blocks get evicted and every read pays an extra seek before it even reaches data.
  2. L2 from the hot data set. Estimate the bytes of data blocks touched in a typical hour; the method is in HBase Read Performance, in depth. Then size L2 above that, because block granularity inflates it: a 1 KB row in a 64 KB block occupies 64 KB of cache.
  3. Check the budget. MemStore plus L1 within the heap rule, L2 plus buffers within direct memory, and several gigabytes left for the OS, the DataNode and the page cache, which still absorbs some L2 misses.
  4. Leave MemStore alone unless writes need it. Taking memory from MemStore to give to cache causes more flushes, more small files and more compaction, and more files means more index and bloom blocks per read.

Inside the on-heap LRU: priorities and factors

The on-heap LruBlockCache is not a single LRU list. Each block has a priority: single on first access, multi once read again, and memory for column families marked IN_MEMORY (the catalog tables use it). When the cache grows past the acceptable factor, eviction frees space until it is back to the minimum factor, taking first from the priority groups that are over their share.

SettingDefaultMeaning
hbase.lru.blockcache.single.percentage0.25Share for blocks read once
hbase.lru.blockcache.multi.percentage0.50Share for blocks read more than once
hbase.lru.blockcache.memory.percentage0.25Share for IN_MEMORY families
hbase.lru.blockcache.acceptable.factor0.99Eviction starts above this fraction of capacity
hbase.lru.blockcache.min.factor0.95Eviction frees down to this fraction
hbase.lru.blockcache.hard.capacity.limit.factor1.2Above this multiple of the acceptable size, inserts are refused
hfile.block.cache.policyLRUL1 policy: LRU or TinyLFU

The split protects the cache from scans: a large one-pass read fills only the single-access share, and blocks read repeatedly survive in the multi share. The memory share is protected space, so marking a big family IN_MEMORY because it is important takes cache from everyone else. Use it only for small, hot lookup families. When inserts arrive faster than eviction can free space, the cache refuses them and counts each refusal in blockCacheFailedInsertionCount; a rising count means the read still succeeded, but the block was not cached.

BucketCache: bucket sizes and the 17 KB trap

The BucketCache divides its space into buckets of fixed size classes and stores each block in the smallest class that fits. The default classes run from 5 KB to 513 KB; each is a power-of-two-ish size plus 1 KB of headroom, for example 17 KB, 33 KB and 65 KB. That headroom matters because HFile blocks are not exact. The writer closes a block once it passes BLOCKSIZE, so a block is the configured size plus part of its last cell, plus a header.

Here is the trap. A family with a 16 KB BLOCKSIZE and cells of a few kilobytes produces many blocks a little over 17 KB. Each goes into a 33 KB bucket, and nearly half of L2 is wasted. The fix is hbase.bucketcache.bucket.sizes: a comma-separated list, smallest first, where every size must be a multiple of 256 bytes or reads fail with an invalid block magic error. Add classes that match your real block sizes, such as 20 KB and 24 KB, after measuring them with the HFile pretty-printer.

<!-- hbase-site.xml: 64 GB off-heap L2, small L1, bucket classes for ~16-20 KB blocks -->
<property><name>hfile.block.cache.size</name><value>0.15</value></property>
<property><name>hbase.bucketcache.ioengine</name><value>offheap</value></property>
<property><name>hbase.bucketcache.size</name><value>65536</value></property>
<property><name>hbase.bucketcache.bucket.sizes</name>
  <value>5120,9216,17408,20480,24576,33792,66560,132096,263168,525312</value></property>

Each value above is a multiple of 256, and the largest class must still hold your largest block, or that block is never cached. Two more settings are worth knowing. Blocks reach the BucketCache through writer queues, by default 3 threads with 64 entries each, and if a queue is full the block is dropped rather than making the read wait. A file-backed engine can keep its contents across restarts if you set hbase.bucketcache.persistent.path, which removes the cold-cache period after a rolling restart.

Per-table and per-request controls

Server settings size the cache; table and request settings decide what deserves to be in it.

# HBase shell: per-family cache behaviour
alter 'events', {NAME => 'raw', BLOCKCACHE => 'false'}             # write-mostly, read by batch scans
alter 'profiles', {NAME => 'p', BLOCKSIZE => '16384',
                   DATA_BLOCK_ENCODING => 'FAST_DIFF',
                   PREFETCH_BLOCKS_ON_OPEN => 'true'}              # point reads, warm on region open
alter 'geo_lookup', {NAME => 'g', IN_MEMORY => 'true'}             # small and read constantly

# Per request: keep a full scan from flushing the cache
scan 'profiles', {CACHE_BLOCKS => false, LIMIT => 10}
  • Scans. In Java, scan.setCacheBlocks(false) for any scan that reads data once, including MapReduce and Spark jobs over HBase. One unflagged nightly export can evict the whole working set.
  • BLOCKCACHE false for families that are written heavily and read only in bulk. Index and bloom blocks are still cached; only data blocks are skipped.
  • Prefetch loads a file's blocks in the background when it opens, so a moved region is warm before traffic finds it. It costs a burst of disk reads, so use it only for latency-critical families that fit in cache.
  • Compression in cache. hbase.block.data.cachecompressed (default false) keeps data blocks in their compressed form, fitting more of them per gigabyte at the cost of decompressing on every hit. It suits a hot set slightly larger than L2 and spare CPU.
  • Writes and compactions. hbase.rs.cacheblocksonwrite (default false) caches blocks as flushes write them, useful when recent data is read immediately. hbase.rs.cachecompactedblocksonwrite does the same for compaction output, so a major compaction does not leave the hot rows cold.

Measuring: the metrics that matter

Tune from metrics, and compare them before and after each change. Every RegionServer exposes them over JMX on its info port (16030 by default) under Hadoop:service=HBase,name=RegionServer,sub=Server.

import time, requests

BEAN = "Hadoop:service=HBase,name=RegionServer,sub=Server"
KEYS = ["blockCacheHitCount", "blockCacheMissCount", "blockCacheEvictionCount",
        "blockCacheFailedInsertionCount", "l1CacheHitCount", "l1CacheMissCount",
        "l2CacheHitCount", "l2CacheMissCount"]

def sample(host):
    bean = requests.get(f"http://{host}:16030/jmx", params={"qry": BEAN}, timeout=5).json()["beans"][0]
    return {k: bean[k] for k in KEYS} | {"express": bean["blockCacheExpressHitPercent"],
                                         "staticIndex": bean["staticIndexSize"], "staticBloom": bean["staticBloomSize"]}

def report(host, interval=60):
    a = sample(host); time.sleep(interval); b = sample(host)
    d = {k: b[k] - a[k] for k in KEYS}
    ratio = lambda h, m: h / max(1, h + m)
    print(f"{host}: hit {ratio(d['blockCacheHitCount'], d['blockCacheMissCount']):.1%} "
          f"L1 {ratio(d['l1CacheHitCount'], d['l1CacheMissCount']):.1%} "
          f"L2 {ratio(d['l2CacheHitCount'], d['l2CacheMissCount']):.1%} "
          f"evict/s {d['blockCacheEvictionCount'] / interval:.0f} "
          f"failed/s {d['blockCacheFailedInsertionCount'] / interval:.0f} "
          f"index+bloom {(b['staticIndex'] + b['staticBloom']) / 2**30:.1f} GiB")

Read the numbers in this order. Use interval deltas, not lifetime ratios; a counter since startup hides last hour's regression. Prefer the express hit percent (blockCacheExpressHitPercent) to the plain one, because it counts only requests that asked to be cached, so flagged scans do not distort it. L1 should be close to 100%; if not, it is too small for the index and bloom footprint. A steady eviction rate with a low L2 hit ratio means the hot set does not fit. The per-type miss counters, such as blockCacheDataMissCount and blockCacheLeafIndexMissCount, show which block type is missing.

Worked example: a profile service

A profile service reads 1 KB rows by key from one table. Each RegionServer has 128 GB of RAM, a 31 GB heap with the default 0.4 and 0.4 split, and no BucketCache. The numbers are illustrative. Over a busy hour the script shows a 58% hit ratio, about 2,500 evictions per second, and old-generation GC activity rising with traffic. The hot set, estimated from key access logs, is about 45 GB per server, against 12.4 GB of L1, and staticIndexSize plus staticBloomSize come to 1.8 GB.

The change: L1 down to 0.15 of heap, which is 4.65 GB and comfortably above the 1.8 GB index and bloom footprint; MemStore unchanged at 0.4, so the heap rule holds at 0.55; an off-heap BucketCache of 64 GB, with direct memory raised to cover it plus buffers; and custom bucket classes, after block statistics showed most data blocks between 17 and 20 KB. The nightly export job also gets setCacheBlocks(false). After one day the L1 hit ratio is near 100%, L2 serves most data reads, evictions drop to a trickle and the old-generation pressure goes away, because data blocks no longer live on the heap. Remaining pause sources are covered in HBase JVM GC Tuning in Depth.

Failure modes

SymptomLikely causeFix
RegionServer will not start after a config changeMemStore plus L1 breaks the heap ruleLower one fraction; keep the sum at 0.8 or below
OutOfMemoryError: Direct buffer memoryL2 plus buffers exceed the direct memory limitRaise the limit or shrink L2
Hit ratio collapses nightlyBatch scans caching blockssetCacheBlocks(false) on bulk readers
L2 half empty but blocks keep missingBlock sizes just above a bucket classAdd matching bucket sizes, multiples of 256
Latency spikes after every restartCold cachePrefetch on hot families, or a persistent file-backed L2
L1 misses on index blocksL1 smaller than the index and bloom footprintSize L1 from staticIndexSize and staticBloomSize

What to do next

  1. Run the JMX script against every RegionServer for an hour of normal traffic and record hit ratios, eviction rate, failed insertions and the index and bloom footprint.
  2. Compare the footprint with L1. If L1 is not at least 1.5 times larger, fix that first.
  3. Estimate the hot data set per server and compare it with L2. If it does not fit, plan an off-heap or file-backed BucketCache, and work out direct memory before changing anything.
  4. Find every bulk reader of HBase, including Spark and MapReduce jobs, and make sure it scans with block caching off.
  5. Inspect real block sizes for your main families and adjust the bucket classes if blocks cluster just above a class.
  6. Change one RegionServer, compare it with its neighbours for a full daily cycle, then roll out and keep the script in your dashboards.
Key takeaway: Size L1 from the index and bloom footprint and L2 from the hot data set, keep MemStore plus L1 within the heap rule, and put large data caches off-heap so they stay out of the garbage collector's way. Match bucket classes to real block sizes, keep bulk scans out of the cache, and use IN_MEMORY sparingly. Measure interval hit ratios, evictions and failed insertions before and after every change, one RegionServer at a time.