HBase has four data operations: Put, Get, Scan and Delete. Their client calls fit on one slide, but their behaviour does not. A Delete writes data rather than removing it. A Put with an old timestamp can be invisible the moment it succeeds. A Get is a Scan in disguise. And a Scan is a stateful conversation with a server, which can time out, return half a row, or evict your hottest cache blocks.
This article explains what each operation does inside a RegionServer: the write path, the WAL and the MemStore, MVCC read points, timestamps and versions, delete markers and compaction, and the scan RPC protocol with its caching, size and lease settings. A worked device-telemetry table ties it together. For the Java client surface itself (Connection, Table, the three delete methods, Result), see the HBase Java API. This page assumes HBase 2.x.
One storage model under four operations
Everything HBase stores is a cell, addressed by row key, column family, qualifier and timestamp, and tagged with a type: Put or one of the delete marker types. Within a region, each column family is a store made of one MemStore and a set of immutable, sorted HFiles. Cells sort by row, then family and qualifier, then timestamp descending, so the newest version comes first.
From that model, the four operations reduce to two primitives. Mutations (Put and Delete) append cells. Reads (Get and Scan) merge every place a cell might live and apply rules about versions, markers and time ranges. Internally a Get is a Scan whose start and stop row are the same key. Once you see it that way, most HBase surprises stop being surprising.
The write path: Put and Delete
The client looks up which region owns the row (through hbase:meta, cached) and sends the mutation to that RegionServer. The server then does the following:
- Takes the row lock, so mutations to one row are serialised and the row is updated atomically, even across families.
- Obtains an MVCC write number and assigns timestamps. Any cell sent without an explicit timestamp gets the server's current time in milliseconds.
- Appends the edit to the write-ahead log and syncs it according to the mutation's
Durability:USE_DEFAULT(the table setting, normallySYNC_WAL),ASYNC_WAL,FSYNC_WALorSKIP_WAL. - Inserts the cells into the MemStore and releases the lock.
- Advances the read point once all earlier write numbers have finished, which is the moment readers can see the write.
SKIP_WAL is faster, and it loses every unflushed edit when a RegionServer dies. Use it only for data you can regenerate. Once a MemStore reaches its flush size, it is written out as a new HFile; compactions later merge HFiles. MemStore and WAL covers flush and log-roll tuning. Batched and buffered writes follow the same per-row path; batching operations covers partial failure.
Timestamps and versions
A family keeps up to VERSIONS versions of each column. The default has been 1 since HBase 0.96. Older versions become invisible to reads at once and are removed physically at compaction. TTL expires cells by age, and MIN_VERSIONS keeps a floor of versions even past TTL.
Client-supplied timestamps are allowed, and they are the source of most "my write disappeared" tickets. If a Put carries a timestamp older than the newest version and VERSIONS => 1, it succeeds, and a Get still returns the newer cell. Two writers whose clocks are 200 ms apart will fight over which value wins. Use explicit timestamps only when the timestamp is the data, such as an event time used for idempotent replays, and document that the larger timestamp wins, not the later write.
Deletes are markers
A Delete takes the same path as a Put and appends a marker cell. There are four marker types. Delete masks one version of a column. DeleteColumn masks all versions of a column with a timestamp at or below the marker's. DeleteFamily masks every column in a family at or below the timestamp. DeleteFamilyVersion masks every column's version at one exact timestamp. Reads apply markers while merging. The masked cells and the markers stay on disk until a major compaction rewrites the store without them.
Three consequences follow:
- Deletes cost a write, and the space comes back late. A table that deletes heavily grows until major compaction runs, and scans slow down while they skip thousands of masked cells. See HBase compaction.
- Deleting the latest version needs a read. Deleting a single version without a timestamp makes the server look up the newest version's timestamp first. That makes it measurably slower than deleting all versions of the column.
- Deletes can mask future puts. Under the classic rules, a
DeleteColumnat time T hides any Put with a timestamp at or below T, including Puts written after the Delete, until a major compaction removes the marker. Delete a row, then re-insert it with an event-time timestamp from yesterday, and the insert is invisible. HBase 2.0 added the family attributeNEW_VERSION_BEHAVIOR, which applies write order (MVCC) so that a later Put is not masked. It has a cost, so test it before enabling it widely.
KEEP_DELETED_CELLS => TRUE keeps masked cells for time-range reads that predate the delete, which is useful for audit and point-in-time queries. Space then comes back only when VERSIONS or TTL retire the cells.
The read path: Get as a one-row scan
When a Get or Scan opens on a region, the server fixes the current MVCC read point. The read sees every write that completed before that point and nothing after it, so a single row is always consistent. For each store, a StoreScanner merges the MemStore with the HFiles through a min-heap ordered by key. It skips HFiles whose time range or Bloom filter rules them out, reads blocks through the block cache, applies delete markers, then trims to the requested versions and TTL. Filters run after that, on the server.
This is why read cost depends on the number of HFiles as well as on the data returned. A Get on a store with 30 HFiles and no Bloom filter may touch 30 blocks. Setting setTimeRange on a Get lets the server skip HFiles outside the range, and narrowing to the column you need avoids reading other families at all.
The scan protocol: caching, size limits and leases
A Scan is a series of RPCs. The first opens a server-side scanner and returns a scanner ID. Each later next call returns a batch of rows, and a final call closes the scanner. Your ResultScanner hides that loop, but these settings control it:
| Setting | What it bounds | Default in 2.x |
|---|---|---|
setCaching(n) | rows per RPC | hbase.client.scanner.caching, unlimited by default, so size is the real limit |
setMaxResultSize(b) | bytes per RPC | hbase.client.scanner.max.result.size, 2 MB |
setBatch(n) | cells per Result; splits wide rows | off |
setAllowPartialResults(true) | lets one row arrive in pieces | false: the client stitches rows |
setLimit(n) | total rows, then the scan closes server-side | none |
setCacheBlocks(false) | keeps full scans out of the block cache | true |
setReadType(...) | PREAD for short scans, STREAM for long ones | DEFAULT switches adaptively |
The server holds a lease on each open scanner. If the client waits longer than hbase.client.scanner.timeout.period (60 seconds by default) between calls, for example because it does slow per-row work inside the loop, the lease expires and the next call fails with a scanner timeout. The client may reopen the scan from the last row, or it may throw, depending on the error. Do heavy work after the scan, or hand rows to a queue. When a filter skips huge numbers of rows, the server sends empty heartbeat responses so the lease stays alive. Scans give per-row consistency, not a snapshot of the table: rows read later in a long scan can reflect writes made after the scan started. Use server-side filters to shrink what crosses the network, but remember that a filter still reads every row in the range.
Worked example: a device telemetry table
Consider a telemetry table holding the last 30 days of readings for two million devices. Queries are "latest reading for device X" and "device X's readings in a time window". The row key is the device ID plus a reversed timestamp, so the newest row for a device sorts first. A short hash prefix spreads sequential device IDs across regions (see hotspotting):
create 'telemetry', {NAME => 'r', VERSIONS => 1, TTL => 2592000,
BLOOMFILTER => 'ROW', COMPRESSION => 'ZSTD'}static byte[] key(String device, long tsMillis) {
byte[] d = Bytes.toBytes(device);
byte salt = (byte) Math.floorMod(device.hashCode(), 16);
return Bytes.add(new byte[]{salt}, d, Bytes.toBytes(Long.MAX_VALUE - tsMillis));
}
static byte[] prefix(String device) {
return Bytes.add(new byte[]{(byte) Math.floorMod(device.hashCode(), 16)}, Bytes.toBytes(device));
}
// Write: event time lives in the key, cell timestamp left to the server.
Put put = new Put(key(dev, eventTs))
.addColumn(R, Bytes.toBytes("temp"), Bytes.toBytes(21.5))
.addColumn(R, Bytes.toBytes("bat"), Bytes.toBytes(87));
table.put(put);
// Latest reading: a prefix scan with limit 1 (newest sorts first).
Scan latest = new Scan().setStartStopRowForPrefixScan(prefix(dev)).setLimit(1);
// Window [from, to): reversed timestamps flip the bounds.
Scan window = new Scan()
.withStartRow(key(dev, to - 1))
.withStopRow(key(dev, from), true)
.addFamily(R)
.setMaxResultSize(4L << 20)
.setCacheBlocks(true);
try (ResultScanner rs = table.getScanner(window)) {
for (Result r : rs) { handle(r); } // keep handle() fast: lease is 60 s
}
// Erase one bad reading: delete the row's family.
table.delete(new Delete(key(dev, badTs)).addFamily(R));Event time lives in the row key rather than the cell timestamp, so a late-arriving reading is a new row and cannot be hidden by a delete marker or by version limits. The TTL expires old rows without any delete traffic. In the reversed key, to - 1 sorts before from, so the window scan runs forward. setStartStopRowForPrefixScan arrived in HBase 2.5, replacing the deprecated setRowPrefixFilter; on older clients, compute the stop row yourself. The prefix must include the device's full ID plus a terminator, or a fixed-length ID, otherwise device ab1 also matches ab12. For a nightly export, use a separate scan with setCacheBlocks(false) and STREAM reads, so analytics do not evict the dashboard's working set.
Failure modes
- Invisible writes. A Put with an explicit timestamp below a newer version, or below a delete marker. Check with
get 't', 'row', {RAW => true, VERSIONS => 10}in the shell, which shows markers. ScannerTimeoutExceptionorUnknownScannerException. Slow processing betweennext()calls, or a RegionServer restart. Shrink the work per batch, or raise the timeout on both the client and the server.- Scans slowing over weeks. Delete markers and expired cells piling up between major compactions, or too many HFiles per store.
- Cache churn and latency spikes. Full-table scans with block caching left on.
- RegionServer memory pressure. Huge rows returned whole. Set
setBatchor allow partial results, and fix the schema if one row holds millions of cells. - Data loss after a crash.
SKIP_WALorASYNC_WALused on data that is not reproducible.
Trade-offs
Single-row atomicity is the only transaction HBase offers. checkAndMutate, increment and append build on the row lock, while multi-row consistency is your problem. Explicit timestamps give idempotent replays, at the price of the masking rules above. Many versions per cell give cheap history, but make every read merge more data. Deletes are cheap to issue and expensive to reclaim, so a TTL is almost always better than a delete job for age-based retention. Large scan caching cuts RPCs but increases memory per call and the time between lease renewals. Set all of these from measured row sizes, not defaults.
What to do next
- Look at each table's families with
describe, and write downVERSIONS,TTL,KEEP_DELETED_CELLSand Bloom filter settings next to the queries they serve. - Search your code for explicit timestamps on Puts and Deletes; justify each one or remove it.
- Replace age-based delete jobs with TTL, and confirm that major compactions run on a schedule you know.
- For every long scan, set
setMaxResultSize, decide onsetCacheBlocks, and move slow processing out of the iteration loop. - Use
RAW => truegets in a staging shell to see markers and versions for a row that behaves oddly. - If you rely on delete-then-reinsert with old timestamps, test
NEW_VERSION_BEHAVIORon a copy of the table. - Confirm the durability level of every write path matches how replaceable its data is.