Few HBase settings are misread as often as IN_MEMORY. The name suggests a column family whose data lives in RAM, like an in-memory database table. It is nothing of the sort. An IN_MEMORY family is written, flushed and stored on HDFS exactly like any other; the flag only changes the priority its blocks get once they are in the block cache. Used on the right family it keeps small, hot reference data cached through scans and churn. Used on the wrong one it does nothing, or in one cluster-wide mode it can starve every other table of cache.
This article explains the block cache's three priorities, the eviction algorithm as it is written in the HBase 2.6 source, how the flag behaves with BucketCache, how it differs from similarly named settings, then sizes a worked example and lists the failure modes. Configuration names and defaults were read from the HBase branch-2.6 source on 2026-10-03.
The read path in one picture
What IN_MEMORY is not
Writes to an IN_MEMORY family follow the normal path: the WAL, the memstore, a flush to a new HFile on HDFS, then compactions. Nothing is loaded into memory when a region opens unless you ask for prefetching separately, and a block reaches the cache only when something reads it. If the region server restarts, the cache starts cold like any other. What the flag changes is a single boolean passed when a block is cached: blocks from that family are cached at the memory priority instead of the usual single-access priority.
That makes IN_MEMORY a cache-policy hint, the HBase relative of marking a table as sticky. It is useful precisely when a family is small and read constantly, and in competition for cache with large tables whose scans would otherwise push its blocks out.
Block cache priorities
The on-heap LruBlockCache sorts every cached block into one of three priorities. A block read for the first time enters as single. If it is read again while cached it is promoted to multi. A block from an IN_MEMORY family enters as memory and stays there; it is never promoted or demoted. Each priority has a target share of the cache set by three factors, whose defaults in the source are:
| Setting | Default | Meaning |
|---|---|---|
hbase.lru.blockcache.single.percentage | 0.25 | share for blocks read once |
hbase.lru.blockcache.multi.percentage | 0.50 | share for blocks read more than once |
hbase.lru.blockcache.memory.percentage | 0.25 | share for IN_MEMORY family blocks |
hbase.lru.rs.inmemoryforcemode | false | cluster-wide switch that protects memory blocks absolutely |
hfile.block.cache.size | 0.4 | fraction of heap given to the L1 block cache |
The single and multi split is what makes the cache scan resistant: a large scan fills the single partition with blocks read once, and eviction takes from there before it touches blocks that proved useful by being read twice. The memory partition sits beside them with its own budget.
How eviction treats each priority
Eviction runs when the cache grows past its acceptable size, 99 percent of capacity by default, and frees blocks until it is back to its minimum size, 95 percent. Each priority's target size is its factor times capacity times the 0.95 minimum factor. In the default mode the algorithm looks at how far each partition is over its target, its overflow, and frees least-recently-used blocks from the partitions that overflow, sharing the work among them. A partition under its target loses nothing.
Two consequences follow. First, an unused share is not wasted: if no IN_MEMORY family is being read, the memory partition is empty, and single and multi blocks can grow past their own targets until the cache is full; eviction then takes from them. Second, the memory partition is protected up to its share, not absolutely. If IN_MEMORY families together have more hot data than about 24 percent of the cache, the memory partition overflows and evicts its own least-recently-used blocks, while the other partitions are left alone.
Force mode changes this. With hbase.lru.rs.inmemoryforcemode=true, or a memory factor of 1.0, eviction empties the single and multi partitions before it evicts any memory block, and when it does not need to touch memory blocks it keeps single and multi at a 1:2 ratio. That is the behaviour people imagine IN_MEMORY always has: memory blocks evicted only after everything else. It is a cluster-wide setting on each region server, and it is dangerous for exactly that reason, as the failure modes below show.
# Simplified from LruBlockCache.evict() in HBase 2.6
bytes_to_free = size - capacity * min_factor # down to 95%
if force_in_memory or memory_factor > 0.999:
if bytes_to_free > single.size + multi.size:
free(single, all); free(multi, all) # memory blocks go last
free(memory, bytes_to_free - freed)
else:
rebalance single and multi to 1:2, never touch memory
else:
for bucket in sorted([single, multi, memory], key=overflow):
share = (bytes_to_free - freed) / buckets_left
free(bucket, min(bucket.overflow(), share)) # only over-target buckets lose blocks
IN_MEMORY with BucketCache
Most production clusters run an off-heap BucketCache as a second level. In HBase 2.x the CombinedBlockCache routes blocks by type: index and bloom blocks go to the on-heap LRU cache (L1), and data blocks go to the BucketCache (L2). The inMemory flag travels with the block to whichever cache gets it.
So with BucketCache enabled, an IN_MEMORY family's data blocks compete inside the bucket cache, whose own partitions are set by hbase.bucketcache.single.factor, hbase.bucketcache.multi.factor and hbase.bucketcache.memory.factor, with the same 0.25, 0.50 and 0.25 defaults, but its own minimum factor of 0.85 rather than 0.95. The LRU percentages above then govern only the family's index and bloom blocks. Sizing the flag means sizing against the right cache: the L2 capacity for data, the L1 capacity for index and bloom blocks. The force-mode switch belongs to the LRU cache only.
Neighbouring settings
Several settings sound similar and do different jobs.
IN_MEMORYsets cache priority for the family's blocks. Default false.IN_MEMORY_COMPACTIONcontrols in-memory compaction inside the memstore, a write-path feature that flattens and merges memstore segments before a flush. It has nothing to do with the block cache.PREFETCH_BLOCKS_ON_OPENreads a store file's blocks into the cache when the file is opened, so the cache warms after a region opens, moves or a compaction writes a new file. Default false. It pairs well with IN_MEMORY for small reference families, since it fills the protected partition without waiting for user reads.BLOCKCACHE => falsestops data blocks of the family from being cached at all, the opposite tool, used for families that are scanned once and never re-read.
HBase itself uses the flag where it fits best: the catalog table hbase:meta is created with its families set IN_MEMORY, with small blocks and row-and-column bloom filters, because every client region lookup that misses its own cache reads it.
Enabling it
Set the flag per family in the shell or the Java API. Altering a family reopens the table's regions to apply the new descriptor.
# HBase shell
create 'merchant_risk', {NAME => 'r', IN_MEMORY => 'true', BLOCKSIZE => '16384',
BLOOMFILTER => 'ROW', PREFETCH_BLOCKS_ON_OPEN => 'true'}
alter 'merchant_risk', {NAME => 'r', IN_MEMORY => 'true'}
describe 'merchant_risk'// Java client, HBase 2.x
ColumnFamilyDescriptor cf = ColumnFamilyDescriptorBuilder
.newBuilder(Bytes.toBytes("r"))
.setInMemory(true)
.setBlocksize(16 * 1024)
.setBloomFilterType(BloomType.ROW)
.setPrefetchBlocksOnOpen(true)
.build();
admin.modifyColumnFamily(TableName.valueOf("merchant_risk"), cf);The priority is recorded when a block is cached, so it is a reasonable inference from the code that blocks cached before the change keep their old priority until they are evicted or the regions reopen; plan to measure after the cache has turned over, not immediately after the alter.
Worked example: merchant risk lookups
A payments platform scores every card transaction. Each score does point Gets against a merchant_risk table of 2 million merchants at about 1 KB per row, roughly 2 GB of data, and against a large transactions table that analysts also scan for reports. Twelve region servers run 32 GB heaps with the default 0.4 block cache and a 48 GB off-heap BucketCache. During report scans, merchant Gets slow down from cache misses.
Size the data. 2 GB across 12 servers is about 170 MB per server if regions are balanced; allow some headroom for index and uneven distribution and call it 250 MB. Size the partition. With BucketCache on, merchant data blocks live in L2: 48 GB times 0.25 times BucketCache's own minimum factor of 0.85 gives roughly 10 GB of memory partition per server, which holds the merchant family many times over. Index and bloom blocks go to L1, where the memory partition is 12.8 GB times 0.25 times 0.95, about 3 GB. Both fit comfortably, so IN_MEMORY plus prefetch on open should keep merchant reads cached through scans. Without BucketCache, the same arithmetic on the 12.8 GB L1 gives about 3 GB for everything IN_MEMORY, still enough here.
Fix the scans too. The report scans should set setCacheBlocks(false) so they stop flushing everyone's single partition. Measure. Compare region server blockCacheHitCount, blockCacheMissCount, blockCacheExpressHitPercent and blockCacheEvictionCount before and after, and the client-side p99 of merchant Gets during a report window. If the p99 does not move, the misses were somewhere else, often index blocks or a hot region.
Check the diagnosis first. A merchant block read twice is already multi priority, and a scan alone fills only the single partition, so a pure scan should not evict it unless the multi partition had grown past its 50 percent target by borrowing idle space. If merchant blocks are being lost, something is crowding the multi partition: here, point lookups on recent transactions re-read their own blocks and compete for the same 50 percent share. IN_MEMORY helps because it moves the merchant family out of that contest into a partition nothing else uses. If instead the multi partition is mostly merchant data already, the flag will change little and the real fix is a larger cache or smaller blocks. That is why the before-and-after measurement matters more than the setting itself.
Failure modes
- Flagging a large family. A 500 GB family marked IN_MEMORY overflows its partition and thrashes inside it, so it gains nothing and the genuinely hot blocks in it lose their protection from each other.
- Flagging everything. If every family is IN_MEMORY, every block is in one partition and the scan resistance of the single and multi split disappears.
- Force mode with a big in-memory set. With force mode on, a large IN_MEMORY family evicts every single and multi block on the server before losing any of its own, and the hit ratio of every other table collapses.
- Expecting a warm restart. The cache is empty after a restart or region move. Use prefetch on open, or a persistent BucketCache where your version supports it, rather than assuming IN_MEMORY persists.
- Confusing the settings. Setting IN_MEMORY when you meant IN_MEMORY_COMPACTION, or the reverse, changes nothing you were trying to change.
- Sizing against the wrong cache. With BucketCache, data blocks are not in the heap cache; tuning LRU percentages will not help a data-block miss problem.
Trade-offs
IN_MEMORY spends a fixed slice of cache on a chosen family. That is a good trade when the family is small, hot and latency-critical, and the competing traffic is large and less valuable per byte. It is a bad trade when the family is large, or when the whole working set already fits, since the partition borrowing behaviour means you gain little. Alternatives include bloom filters and smaller block sizes to cut the bytes read per Get, region replicas to spread read load, a separate table on a dedicated RegionServer group, or an application cache in front of HBase when the data changes rarely.
Related reading on this site: the HBase block cache, BucketCache, column family design, bloom filters and the hbase:meta table.
What to do next
- List families with small size and heavy point reads; measure their size per region server.
- Check whether BucketCache is on, and size the memory partition of the cache that will actually hold the data blocks.
- Set IN_MEMORY only where the family fits comfortably in that partition, and add PREFETCH_BLOCKS_ON_OPEN for reference data.
- Make large analytic scans use setCacheBlocks(false).
- Leave hbase.lru.rs.inmemoryforcemode off unless you have proved the in-memory set is small and bounded.
- Record hit, miss, eviction and p99 latency before and after, during the busiest scan window.
- Review the flagged families each quarter, as table sizes grow.