A column family in HBase is often introduced as a way to group columns. That description is true and nearly useless. A family is the unit at which HBase decides how many versions of a cell to keep, how long cells live, what a delete hides, how bytes are encoded on disk and which files a read has to open. Two cells with the same row and qualifier but in different families behave differently because the family carries those rules.
This article is about those rules at runtime. Whether to have one family or several, and how to choose per-family storage settings, is covered in HBase column family design; this page assumes you have your families and explains what they do to every write and read: the cell format, versions and time to live, the delete markers and their surprising interactions, the read path, and how to inspect what a family actually holds.
A cell, byte by byte
HBase stores everything as cells, and the coordinates of a cell are row, family, qualifier and timestamp. On disk each cell is a key-value record. The key holds a row length and the row bytes, a family length and the family bytes, the qualifier, an eight-byte timestamp and a one-byte type that says whether the cell is a put or one of the delete markers. Two four-byte length fields for key and value come first, and HFile version 3 adds a tag section. The family name is therefore repeated in every cell, which is why short family names matter and why the design article recommends one-letter names.
The fixed overhead is easy to compute. With a 16-byte row key, family d, an 8-byte qualifier and a 10-byte value, the record is 4 + 4 + 2 + 16 + 1 + 1 + 8 + 8 + 1 + 10 = 55 bytes, of which only 10 are the value. Block encoding such as FAST_DIFF removes much of the repeated row prefix inside a block, and compression shrinks the rest, but the block cache holds decoded or encoded blocks depending on configuration, so overhead per cell still shapes how much of a family fits in memory.
Versions: what VERSIONS really bounds
Each family has a VERSIONS setting, the maximum number of timestamped versions kept per row and column. The default has been 1 since HBase 0.96. A put with a new timestamp does not overwrite the old cell in place; it adds a cell, and older versions beyond the limit become invisible to reads immediately and are physically removed when a compaction rewrites the files. Until then they still occupy disk and are still read and skipped by scanners.
Reads ask for versions explicitly. A Get returns one version by default; Get.readVersions(n) in the 2.x client asks for up to n, and setTimeRange restricts by timestamp. Asking for more versions than the family keeps returns at most the family limit, because versions above it are already filtered. Timestamps are client-settable, which is powerful and dangerous: a client writing with a timestamp from a skewed clock or a fixed constant can produce a newer write that is ranked as older and therefore hidden.
// HBase 2.x client: define a family with explicit version and lifetime rules
ColumnFamilyDescriptor events = ColumnFamilyDescriptorBuilder
.newBuilder(Bytes.toBytes("e"))
.setMaxVersions(5) // keep up to five versions per cell
.setMinVersions(1) // ...and always keep the newest, even past TTL
.setTimeToLive(7 * 24 * 3600) // seconds
.setKeepDeletedCells(KeepDeletedCells.FALSE)
.build();
admin.modifyColumnFamily(TableName.valueOf("profiles"), events);
// Read up to three versions of one column in that family
Get get = new Get(Bytes.toBytes("user#0042"))
.addColumn(Bytes.toBytes("e"), Bytes.toBytes("login"))
.readVersions(3);
Result r = table.get(get);
for (Cell cell : r.getColumnCells(Bytes.toBytes("e"), Bytes.toBytes("login"))) {
System.out.println(cell.getTimestamp() + " " + Bytes.toString(CellUtil.cloneValue(cell)));
}
TTL and MIN_VERSIONS
A family's TTL is measured in seconds against each cell's timestamp. A cell older than the TTL is filtered out of reads as soon as it expires, and it is removed from disk when a compaction rewrites its file; when MIN_VERSIONS is 0 and every cell in an HFile has expired, the whole file can be dropped without being rewritten. TTL is per family, which is one of the legitimate reasons to separate short-lived data such as session events from long-lived profile data.
MIN_VERSIONS interacts with TTL in a way that surprises people. When the two conflict, MIN_VERSIONS wins: with TTL of seven days and MIN_VERSIONS of 1, a column that has not been written for a month still returns its newest value, because HBase keeps at least one version regardless of age. Without MIN_VERSIONS, the same column simply disappears after seven days. Decide which semantics you want, "expire everything old" or "keep the latest, expire history", and set the family accordingly.
Four kinds of delete
A delete in HBase writes a marker, a tombstone, rather than removing data. Markers sort ahead of the cells they cover and hide them during reads until a major compaction removes both. There are four types, and the type you write decides what disappears.
| Marker | Client call | What it hides |
|---|---|---|
| Delete | Delete.addColumn(f, q, ts) | exactly one version of one column, at that timestamp (the latest if none given) |
| DeleteColumn | Delete.addColumns(f, q, ts) | all versions of one column with timestamp up to ts |
| DeleteFamily | Delete.addFamily(f, ts) | all columns of the family in that row, up to ts |
| DeleteFamilyVersion | Delete.addFamilyVersion(f, ts) | all columns of the family in that row at exactly ts |
Note the singular and plural: addColumn removes one version and leaves older versions visible, which is a common source of "the delete did not work" reports on families with VERSIONS above 1. A whole-row new Delete(row) writes a DeleteFamily marker into every family of the table, which is why a row delete costs one marker per family.
Deletes mask puts, and NEW_VERSION_BEHAVIOR
Because markers are matched by timestamp, a marker hides every matching cell, including cells written after it. If you delete a column up to time T and then put a new value with a timestamp at or below T, the new put is invisible. It stays invisible until a major compaction removes the marker, at which point it suddenly reappears. The same mechanism means that results can change after a major compaction for families with several versions, because the compaction drops cells that were being counted towards the version limit.
Two practices avoid this. First, let the server assign timestamps, or make client timestamps strictly monotonic per cell. Second, for families where deletes and re-writes interleave, consider the NEW_VERSION_BEHAVIOR family attribute added in HBase 2.0. It makes the matcher use write order (sequence ids) as well as timestamps, so a put that arrives after a delete is not masked by it and results no longer change across major compactions. It costs some read-path work, so enable it per family where the semantics matter rather than everywhere.
KEEP_DELETED_CELLS and time-travel reads
By default a deleted cell is unreachable. Setting KEEP_DELETED_CELLS on a family changes that for reads with a time range. With TRUE, deleted cells are retained until something else removes them, such as the version limit or TTL, so a Get with a time range ending before the delete still sees the value. With TTL, deleted cells are retained until the delete marker itself expires under the family TTL, which is useful together with MIN_VERSIONS when you want to keep a minimum history but still purge deletes eventually. FALSE is the default.
This is how you build a point-in-time view of a family without a separate audit table: keep several versions, keep deleted cells, and query with setTimeRange(0, t). The price is disk and read cost, because markers and dead cells stay in files that every scan of that family must merge.
The read path, family by family
A region serves a Get or Scan with a region scanner, which opens one store scanner for each family the request needs. Each store scanner builds a heap over the family's MemStore and its HFiles, skips files whose bloom filter or time range rules them out, and merges the rest in key order. A query matcher then applies the family's rules: versions, TTL, MIN_VERSIONS, markers and KEEP_DELETED_CELLS. The region scanner merges the per-family streams into rows.
The consequence is that naming families in a request is a performance tool. Get.addFamily or Scan.addFamily means other families' files are never opened. A scan that names nothing opens every family. When a scan filters on one small family but returns a large one, setLoadColumnFamiliesOnDemand(true) lets the region evaluate the filter on the essential family first and load the other families only for rows that pass; it applies when the filter can report which families are essential, as SingleColumnValueFilter does.
// Filter on the small 'm' family, return the large 'd' family only for matching rows
Scan scan = new Scan()
.addFamily(Bytes.toBytes("m"))
.addFamily(Bytes.toBytes("d"))
.setFilter(new SingleColumnValueFilter(
Bytes.toBytes("m"), Bytes.toBytes("status"),
CompareOperator.EQUAL, Bytes.toBytes("ACTIVE")))
.setLoadColumnFamiliesOnDemand(true)
.setCaching(500);
try (ResultScanner rs = table.getScanner(scan)) {
for (Result row : rs) { process(row); }
}
Inspecting what a family holds
Every store is a directory under the region in HDFS, named after the family, so you can look at a family directly. The HBase shell shows the family attributes; the HFile tool prints file metadata, statistics and, if asked, cells.
# Family attributes as the cluster sees them
hbase shell> describe 'profiles'
hbase shell> alter 'profiles', NAME => 'e', VERSIONS => 5, MIN_VERSIONS => 1, TTL => 604800
# One family's store files for one region (path layout: data/<ns>/<table>/<region>/<family>)
hdfs dfs -ls /hbase/data/default/profiles/<region-encoded-name>/e
# Metadata (-m), key/value statistics (-s) and, sparingly, cells (-p) of one HFile
hbase hfile -m -s -f /hbase/data/default/profiles/<region>/e/<hfile>The statistics report key and value lengths and the number of cells, which lets you check the overhead arithmetic above against reality. Count delete markers by printing cells on a sample file; a family with millions of markers that never sees a major compaction will show them clearly, and so will read latency. On the cluster side, the RegionServer exposes per-store metrics such as store file count and size; a family whose file count keeps climbing is a family whose compactions are falling behind, covered in HBase compaction.
Worked example: a profile table with history
A team stores user profiles in family p (VERSIONS 1, no TTL) and login events in family e (VERSIONS 5, TTL seven days). Support asks why a user who has not logged in for a month shows no last-login time. The answer is the TTL: all five versions expired. Setting MIN_VERSIONS to 1 on e keeps the newest login forever and still expires older history, which is what support wanted.
A month later, GDPR deletions arrive. The deletion job calls Delete.addColumn on e:login and the user's older logins reappear in the UI, because the singular call deleted one version. The fix is addFamily on both families, or a whole-row delete. Then a re-registration bug shows up: the account service writes a new profile with a timestamp from a cached clock that is behind the deletion time, and the new profile is invisible until the weekly major compaction, when it pops back. The team switches to server-assigned timestamps and enables NEW_VERSION_BEHAVIOR on p.
Failure modes
- Deletes that do not delete: singular
addColumnon a multi-version family. UseaddColumnsoraddFamily. - Writes that vanish, then return: client timestamps at or below an existing marker. Use server timestamps or NEW_VERSION_BEHAVIOR.
- Data that should have expired: MIN_VERSIONS greater than zero overriding TTL. Check both settings together.
- Scans slower than the data size suggests: unnamed families, piles of markers or uncompacted versions. Name families and check marker counts and file counts.
- Unexpected point-in-time results: KEEP_DELETED_CELLS set to TRUE without a TTL keeps deleted cells indefinitely, growing every scan of the family.
Trade-offs
Every rule a family carries is a trade between read cost, disk and semantics. More versions and kept deleted cells give history but make every read merge more cells. TTL is the cheapest form of deletion because, with MIN_VERSIONS at 0, expired files can be dropped whole, while explicit deletes add markers that cost until compaction. NEW_VERSION_BEHAVIOR fixes correctness for interleaved deletes at some read cost. Families let you choose these per data set, which is the real reason to have more than one; the storage and block-cache side of that choice is in the HBase block cache and the full schema picture in HBase schema design.
What to do next
- Run
describeon your busiest tables and write down VERSIONS, TTL, MIN_VERSIONS, KEEP_DELETED_CELLS and NEW_VERSION_BEHAVIOR for each family, with the reason for each value. - Search your code for
addColumnin Delete calls and confirm each one is meant to remove a single version. - Find every place a client sets timestamps explicitly and decide whether it should.
- Make reads name their families, and use on-demand family loading for scans that filter on a small family.
- Inspect one HFile per family with
hbase hfile -m -sand compare cell counts and sizes with your expectations. - Check major compaction frequency for families that receive many deletes, and alert on rising store file counts.