Some datasets are mostly holes. A product catalogue knows thousands of attributes, but a T-shirt has fabric and size, not voltage or ISBN. A user-profile store collects hundreds of optional preferences, most never set. In a relational table each missing attribute is a NULL in a fixed column or a row in an entity-attribute-value side table. HBase was designed for this shape: a row can hold any set of columns, and a column that is not there costs nothing.
That is true but incomplete. Cells that do exist are more expensive than most people assume, null has three different meanings, and the default filter behaviour quietly returns rows that lack the column you filtered on. This article builds an accurate cost model, works through a product-attribute table byte by byte, and covers the query, bloom-filter and scanning techniques that make sparse tables fast. It assumes the key-design basics in HBase schema design.
How HBase stores a sparse row
An HBase table declares only its column families. Qualifiers within a family are free-form byte strings chosen at write time, so two rows in the same table can have entirely different columns. Physically a row is not a record; it is a sorted run of cells, and a region stores only cells that were written. An HFile for a table with 4,000 possible qualifiers and 25 present per row contains about 25 cells per row, not 4,000 slots.
This is the sense in which HBase is a sparse store: there is no per-row bitmap, no null marker, no fixed width. Adding a new attribute needs no DDL and no backfill. Readers ask for the qualifiers they care about, or for the whole family, and get back whatever is present.
The cost model: absence is free, presence is not
Absence is free. Presence is not, because each cell is stored as a self-describing KeyValue. In the classic serialization a cell carries: a 4-byte key length, a 4-byte value length, a 2-byte row length, the row key, a 1-byte family length, the family name, the qualifier, an 8-byte timestamp and a 1-byte type, then the value (HFile v3 adds tag bytes when tags are in use). The row key, family and qualifier are repeated in every cell.
Work it through for a catalogue cell. Row key p# plus an 8-character hex hash prefix plus a 6-character product id is 16 bytes; family attr is 4 bytes; qualifier voltage_rating is 14 bytes; the value is a 4-byte integer.
| Component | Bytes (long names) | Bytes (short names) |
|---|---|---|
| Length fields (key, value, row, family) | 11 | 11 |
| Row key | 16 | 16 |
| Family | 4 (attr) | 1 (a) |
| Qualifier | 14 (voltage_rating) | 3 (vlt) |
| Timestamp + type | 9 | 9 |
| Value | 4 | 4 |
| Total | 58 | 44 |
The payload is 4 bytes out of 58. Multiply by 25 cells per product and 50 million products and the raw cell bytes are about 72 GB with long names, of which about 5 GB is data. That ratio is the price of sparseness, and it shows up in four places: memstore size (so flush frequency), RPC payloads, client heap during scans, and, before encoding, block cache and disk.
Two mechanisms recover most of the disk and cache cost. Data block encoding (DATA_BLOCK_ENCODING => 'FAST_DIFF' or ROW_INDEX_V1) stores each key as a difference from the previous one, so the repeated row key and family in consecutive cells of the same row nearly vanish, and encoded blocks stay encoded in the block cache. Compression (COMPRESSION => 'ZSTD' or LZ4) squeezes whole blocks on disk. Neither helps the memstore or the wire, which is why short family names (one character) still matter, and why very short qualifiers are worth considering for high-volume tables. The trade-off is readability: a three-letter qualifier needs a dictionary in code, and that dictionary becomes part of your schema.
Three kinds of empty
A relational column is NULL or not. An HBase cell has three distinct empty states, and applications that conflate them produce subtle bugs:
| State | How it arises | What a Get returns |
|---|---|---|
| Absent | Never written, or deleted and compacted away | No cell for that qualifier |
| Empty value | Put with a zero-length byte array | A cell whose value has length 0 |
| Masked by a delete marker | Delete written but not yet removed by major compaction | No cell, but the marker still occupies space and costs read work |
Pick one meaning and enforce it in a single data-access layer. The usual rule: absent means unknown or not applicable; never write empty values to mean null; to clear an attribute, issue a delete for that column. Know which delete you are issuing: Delete.addColumn removes only the latest version, while Delete.addColumns removes all versions up to the timestamp. Families with many versions and frequent attribute clears accumulate markers until major compaction, so reads over a sparse row can do more work than the visible cells suggest. Version and delete semantics are covered fully in HBase column families.
Worked example: a product attribute catalogue
Design a catalogue table for 50 million products with about 4,000 known attributes, 25 present per product on average, and these reads: fetch one product's attributes; fetch a few named attributes for a page of products; find products where a given attribute has a given value.
Use a single family for attributes: every family is a separate store with its own files and compactions, and families share region-wide flush triggers and splits, so extra families mostly add small files. Hash-prefix the row key to spread writes, keep one version, and encode and compress the family:
create 'catalog', {NAME => 'a', VERSIONS => 1, BLOOMFILTER => 'ROWCOL',
DATA_BLOCK_ENCODING => 'FAST_DIFF', COMPRESSION => 'ZSTD'},
{SPLITS => ['1', '2', '3', '4', '5', '6', '7', '8', '9', 'a', 'b', 'c', 'd', 'e', 'f']}Writes go through one helper that owns the qualifier dictionary and the null rule:
static final byte[] A = Bytes.toBytes("a");
void saveAttributes(Table t, String productId, Map<String, byte[]> attrs) throws IOException {
byte[] row = rowKey(productId); // hash prefix + id
Put put = new Put(row);
Delete clears = new Delete(row);
boolean anyClear = false;
for (Map.Entry<String, byte[]> e : attrs.entrySet()) {
byte[] q = Dict.code(e.getKey()); // "voltage_rating" -> "vlt"
if (e.getValue() == null) { // null means: remove the attribute
clears.addColumns(A, q);
anyClear = true;
} else {
put.addColumn(A, q, e.getValue());
}
}
RowMutations rm = new RowMutations(row); // applied atomically: one row, one RPC
if (!put.isEmpty()) rm.add(put);
if (anyClear) rm.add(clears);
if (!rm.getMutations().isEmpty()) t.mutateRow(rm);
}
Result readSome(Table t, String productId, String... names) throws IOException {
Get g = new Get(rowKey(productId));
for (String n : names) g.addColumn(A, Dict.code(n)); // only the cells you need cross the wire
return t.get(g);
}The third read, attribute equals value, cannot be served by the main table without a full scan, because the row key is the product id. For anything user-facing, maintain an inverted index table keyed by attribute code + value + product id, written in the same helper (or by a coprocessor or a stream job if you need it consistent under failure). For occasional analytics, a filtered scan is acceptable, and that is where the main trap lives.
Querying sparse columns without phantom matches
SingleColumnValueFilter keeps rows that lack the column. By default, a row with no cell for the tested column passes the filter, because there was nothing to compare. On a sparse table that is most rows, so a query for fabric = cotton returns cotton T-shirts and every toaster. Set setFilterIfMissing(true) and, unless you want any version to match, setLatestVersionOnly(true) (the default). Also make sure the scan actually reads the tested column: if you restrict returned columns with addColumn, include the filter column too, or the filter never sees it.
SingleColumnValueFilter f = new SingleColumnValueFilter(
A, Dict.code("fabric"), CompareOperator.EQUAL, Bytes.toBytes("cotton"));
f.setFilterIfMissing(true); // the line that makes sparse queries correct
Scan scan = new Scan()
.addFamily(A) // filter column is inside the scanned family
.setFilter(f)
.setCaching(500);For selecting columns rather than rows, ColumnPrefixFilter and MultipleColumnPrefixFilter return only qualifiers starting with given prefixes, which works well if your dictionary groups related attributes under a shared prefix (dim_ for dimensions, elc_ for electrical). The full filter catalogue and how filters execute on the region server are covered in HBase filters. Remember filters reduce bytes returned, not bytes read: a server-side filter over a full table still scans the full table.
Bloom filters for sparse column reads
Bloom filters let a Get skip HFiles that cannot contain what it wants. The default ROW bloom answers whether a row may be present in a file. A ROWCOL bloom answers whether a specific row and column pair may be present. For sparse tables where reads ask for a few named attributes and a row's cells are spread across several HFiles (because attributes were written at different times), ROWCOL can skip files that hold the row but not the column. It does nothing for scans or for Gets of a whole family, and it is larger because it has one entry per cell, not per row. Use it when the dominant read is a Get with explicit columns, as in the catalogue above; otherwise stay with ROW. Sizing and false-positive tuning are in HBase bloom filters.
When sparse rows become very wide
Sparse tables sometimes grow a few rows with enormous numbers of columns: a vendor that dumps 200,000 attributes onto a test product, or a qualifier-per-event pattern that never stops. HBase never splits a row across regions, so a huge row pins its region and every read of it. By default a Get or a non-batched scan materialises the whole row, and rows above hbase.table.max.rowsize fail with a RowTooBigException.
Scan wide rows in pieces: setBatch(n) caps cells per Result, and setAllowPartialResults(true) lets the client receive part of a row when the result-size limit is reached. Your code must then stitch partial Results that share a row key. Better still, cap columns per row in the write helper and move unbounded collections to a tall layout; the decision is laid out in wide versus tall tables.
Failure modes
- Phantom matches. Value filters without
setFilterIfMissing(true)return rows lacking the column. - Memstore pressure from key overhead. Long qualifiers make a 4-byte value cost more than 10 times its size in memory, so flushes come early and files are small. Shorten names or batch writes.
- Many families for many attribute groups. Each family is a separate store with its own files and compactions, and region-wide flushes under memory or WAL pressure still write small files for sparse families. Prefer one family and qualifier prefixes.
- Empty values as null. Readers that test presence treat an empty cell as set. Delete instead.
- Dictionary drift. Two services encode the same attribute differently. Keep the dictionary in one shared library, append-only, and never reuse a code.
- Unbounded rows. One row grows to gigabytes and its region becomes a hotspot that cannot split.
Trade-offs
Against a relational entity-attribute-value table, HBase avoids the self-joins and gives atomic per-row updates across all attributes, but loses ad hoc querying by value without an index you maintain. Against storing all attributes as one JSON cell per product, per-attribute cells let you read and update single attributes without rewriting the blob, at the cost of per-cell key overhead; a JSON cell is smaller and simpler when attributes are always read and written together. Many teams combine them: hot, individually updated attributes as cells, and the long tail as one serialized blob.
What to do next
- List your reads and decide whether they name columns, whole rows or values; that drives blooms and indexes.
- Measure average cells per row and bytes per cell on a sample with
hbase hfile -m -f <hfile>before and after encoding. - Shorten family names to one character and decide whether a qualifier dictionary is worth it at your volume.
- Enable
FAST_DIFF(orROW_INDEX_V1) and ZSTD or LZ4 on sparse families. - Put all writes behind one helper that enforces the null rule and owns the dictionary.
- Audit every SingleColumnValueFilter for
setFilterIfMissing(true). - Switch to ROWCOL blooms only if the dominant read is a Get with explicit columns, then compare HFiles read per Get.
- Cap columns per row and use
setBatchplus partial results for any row that can grow without bound.