Some datasets are mostly holes. A product catalogue knows thousands of attributes, but a T-shirt has fabric and size, not voltage or ISBN. A user-profile store collects hundreds of optional preferences, most never set. In a relational table each missing attribute is a NULL in a fixed column or a row in an entity-attribute-value side table. HBase was designed for this shape: a row can hold any set of columns, and a column that is not there costs nothing.

That is true but incomplete. Cells that do exist are more expensive than most people assume, null has three different meanings, and the default filter behaviour quietly returns rows that lack the column you filtered on. This article builds an accurate cost model, works through a product-attribute table byte by byte, and covers the query, bloom-filter and scanning techniques that make sparse tables fast. It assumes the key-design basics in HBase schema design.

Advertisement

How HBase stores a sparse row

An HBase table declares only its column families. Qualifiers within a family are free-form byte strings chosen at write time, so two rows in the same table can have entirely different columns. Physically a row is not a record; it is a sorted run of cells, and a region stores only cells that were written. An HFile for a table with 4,000 possible qualifiers and 25 present per row contains about 25 cells per row, not 4,000 slots.

This is the sense in which HBase is a sparse store: there is no per-row bitmap, no null marker, no fixed width. Adding a new attribute needs no DDL and no backfill. Readers ask for the qualifiers they care about, or for the whole family, and get back whatever is present.

Logical view: one table, thousands of possible columnsrow keya:colora:voltagea:isbna:fabric... 4,000 morep#000123redcottonp#000124230VPhysical view: an HFile stores only the cells that exist, sorted by row, family, qualifier, timestampp#000123 / a / color / ts / Put = redp#000123 / a / fabric / ts / Put = cottonp#000124 / a / voltage / ts / Put = 230VEvery stored cell repeats its keyrow + family + qualifier + timestamp + typeEmpty boxes above cost zero bytes. The filled ones cost key overhead plus value.Block encodingFAST_DIFF, ROW_INDEX_V1Block compressionZSTD, LZ4, SNAPPYNot shrunkmemstore, RPC, client heap
Sparse data in HBase. The logical table is a grid with thousands of mostly empty columns; physically only present cells are stored, each carrying its full key. Encoding and compression shrink keys on disk and in the block cache, but not in the memstore or on the wire.

The cost model: absence is free, presence is not

Absence is free. Presence is not, because each cell is stored as a self-describing KeyValue. In the classic serialization a cell carries: a 4-byte key length, a 4-byte value length, a 2-byte row length, the row key, a 1-byte family length, the family name, the qualifier, an 8-byte timestamp and a 1-byte type, then the value (HFile v3 adds tag bytes when tags are in use). The row key, family and qualifier are repeated in every cell.

Work it through for a catalogue cell. Row key p# plus an 8-character hex hash prefix plus a 6-character product id is 16 bytes; family attr is 4 bytes; qualifier voltage_rating is 14 bytes; the value is a 4-byte integer.

ComponentBytes (long names)Bytes (short names)
Length fields (key, value, row, family)1111
Row key1616
Family4 (attr)1 (a)
Qualifier14 (voltage_rating)3 (vlt)
Timestamp + type99
Value44
Total5844

The payload is 4 bytes out of 58. Multiply by 25 cells per product and 50 million products and the raw cell bytes are about 72 GB with long names, of which about 5 GB is data. That ratio is the price of sparseness, and it shows up in four places: memstore size (so flush frequency), RPC payloads, client heap during scans, and, before encoding, block cache and disk.

Two mechanisms recover most of the disk and cache cost. Data block encoding (DATA_BLOCK_ENCODING => 'FAST_DIFF' or ROW_INDEX_V1) stores each key as a difference from the previous one, so the repeated row key and family in consecutive cells of the same row nearly vanish, and encoded blocks stay encoded in the block cache. Compression (COMPRESSION => 'ZSTD' or LZ4) squeezes whole blocks on disk. Neither helps the memstore or the wire, which is why short family names (one character) still matter, and why very short qualifiers are worth considering for high-volume tables. The trade-off is readability: a three-letter qualifier needs a dictionary in code, and that dictionary becomes part of your schema.

Advertisement

Three kinds of empty

A relational column is NULL or not. An HBase cell has three distinct empty states, and applications that conflate them produce subtle bugs:

StateHow it arisesWhat a Get returns
AbsentNever written, or deleted and compacted awayNo cell for that qualifier
Empty valuePut with a zero-length byte arrayA cell whose value has length 0
Masked by a delete markerDelete written but not yet removed by major compactionNo cell, but the marker still occupies space and costs read work

Pick one meaning and enforce it in a single data-access layer. The usual rule: absent means unknown or not applicable; never write empty values to mean null; to clear an attribute, issue a delete for that column. Know which delete you are issuing: Delete.addColumn removes only the latest version, while Delete.addColumns removes all versions up to the timestamp. Families with many versions and frequent attribute clears accumulate markers until major compaction, so reads over a sparse row can do more work than the visible cells suggest. Version and delete semantics are covered fully in HBase column families.

Worked example: a product attribute catalogue

Design a catalogue table for 50 million products with about 4,000 known attributes, 25 present per product on average, and these reads: fetch one product's attributes; fetch a few named attributes for a page of products; find products where a given attribute has a given value.

Use a single family for attributes: every family is a separate store with its own files and compactions, and families share region-wide flush triggers and splits, so extra families mostly add small files. Hash-prefix the row key to spread writes, keep one version, and encode and compress the family:

create 'catalog', {NAME => 'a', VERSIONS => 1, BLOOMFILTER => 'ROWCOL',
                   DATA_BLOCK_ENCODING => 'FAST_DIFF', COMPRESSION => 'ZSTD'},
                  {SPLITS => ['1', '2', '3', '4', '5', '6', '7', '8', '9', 'a', 'b', 'c', 'd', 'e', 'f']}

Writes go through one helper that owns the qualifier dictionary and the null rule:

static final byte[] A = Bytes.toBytes("a");

void saveAttributes(Table t, String productId, Map<String, byte[]> attrs) throws IOException {
    byte[] row = rowKey(productId);                  // hash prefix + id
    Put put = new Put(row);
    Delete clears = new Delete(row);
    boolean anyClear = false;
    for (Map.Entry<String, byte[]> e : attrs.entrySet()) {
        byte[] q = Dict.code(e.getKey());            // "voltage_rating" -> "vlt"
        if (e.getValue() == null) {                  // null means: remove the attribute
            clears.addColumns(A, q);
            anyClear = true;
        } else {
            put.addColumn(A, q, e.getValue());
        }
    }
    RowMutations rm = new RowMutations(row);         // applied atomically: one row, one RPC
    if (!put.isEmpty()) rm.add(put);
    if (anyClear) rm.add(clears);
    if (!rm.getMutations().isEmpty()) t.mutateRow(rm);
}

Result readSome(Table t, String productId, String... names) throws IOException {
    Get g = new Get(rowKey(productId));
    for (String n : names) g.addColumn(A, Dict.code(n));   // only the cells you need cross the wire
    return t.get(g);
}

The third read, attribute equals value, cannot be served by the main table without a full scan, because the row key is the product id. For anything user-facing, maintain an inverted index table keyed by attribute code + value + product id, written in the same helper (or by a coprocessor or a stream job if you need it consistent under failure). For occasional analytics, a filtered scan is acceptable, and that is where the main trap lives.

Querying sparse columns without phantom matches

SingleColumnValueFilter keeps rows that lack the column. By default, a row with no cell for the tested column passes the filter, because there was nothing to compare. On a sparse table that is most rows, so a query for fabric = cotton returns cotton T-shirts and every toaster. Set setFilterIfMissing(true) and, unless you want any version to match, setLatestVersionOnly(true) (the default). Also make sure the scan actually reads the tested column: if you restrict returned columns with addColumn, include the filter column too, or the filter never sees it.

SingleColumnValueFilter f = new SingleColumnValueFilter(
        A, Dict.code("fabric"), CompareOperator.EQUAL, Bytes.toBytes("cotton"));
f.setFilterIfMissing(true);          // the line that makes sparse queries correct

Scan scan = new Scan()
        .addFamily(A)                    // filter column is inside the scanned family
        .setFilter(f)
        .setCaching(500);

For selecting columns rather than rows, ColumnPrefixFilter and MultipleColumnPrefixFilter return only qualifiers starting with given prefixes, which works well if your dictionary groups related attributes under a shared prefix (dim_ for dimensions, elc_ for electrical). The full filter catalogue and how filters execute on the region server are covered in HBase filters. Remember filters reduce bytes returned, not bytes read: a server-side filter over a full table still scans the full table.

Bloom filters for sparse column reads

Bloom filters let a Get skip HFiles that cannot contain what it wants. The default ROW bloom answers whether a row may be present in a file. A ROWCOL bloom answers whether a specific row and column pair may be present. For sparse tables where reads ask for a few named attributes and a row's cells are spread across several HFiles (because attributes were written at different times), ROWCOL can skip files that hold the row but not the column. It does nothing for scans or for Gets of a whole family, and it is larger because it has one entry per cell, not per row. Use it when the dominant read is a Get with explicit columns, as in the catalogue above; otherwise stay with ROW. Sizing and false-positive tuning are in HBase bloom filters.

When sparse rows become very wide

Sparse tables sometimes grow a few rows with enormous numbers of columns: a vendor that dumps 200,000 attributes onto a test product, or a qualifier-per-event pattern that never stops. HBase never splits a row across regions, so a huge row pins its region and every read of it. By default a Get or a non-batched scan materialises the whole row, and rows above hbase.table.max.rowsize fail with a RowTooBigException.

Scan wide rows in pieces: setBatch(n) caps cells per Result, and setAllowPartialResults(true) lets the client receive part of a row when the result-size limit is reached. Your code must then stitch partial Results that share a row key. Better still, cap columns per row in the write helper and move unbounded collections to a tall layout; the decision is laid out in wide versus tall tables.

Failure modes

  • Phantom matches. Value filters without setFilterIfMissing(true) return rows lacking the column.
  • Memstore pressure from key overhead. Long qualifiers make a 4-byte value cost more than 10 times its size in memory, so flushes come early and files are small. Shorten names or batch writes.
  • Many families for many attribute groups. Each family is a separate store with its own files and compactions, and region-wide flushes under memory or WAL pressure still write small files for sparse families. Prefer one family and qualifier prefixes.
  • Empty values as null. Readers that test presence treat an empty cell as set. Delete instead.
  • Dictionary drift. Two services encode the same attribute differently. Keep the dictionary in one shared library, append-only, and never reuse a code.
  • Unbounded rows. One row grows to gigabytes and its region becomes a hotspot that cannot split.

Trade-offs

Against a relational entity-attribute-value table, HBase avoids the self-joins and gives atomic per-row updates across all attributes, but loses ad hoc querying by value without an index you maintain. Against storing all attributes as one JSON cell per product, per-attribute cells let you read and update single attributes without rewriting the blob, at the cost of per-cell key overhead; a JSON cell is smaller and simpler when attributes are always read and written together. Many teams combine them: hot, individually updated attributes as cells, and the long tail as one serialized blob.

What to do next

  1. List your reads and decide whether they name columns, whole rows or values; that drives blooms and indexes.
  2. Measure average cells per row and bytes per cell on a sample with hbase hfile -m -f <hfile> before and after encoding.
  3. Shorten family names to one character and decide whether a qualifier dictionary is worth it at your volume.
  4. Enable FAST_DIFF (or ROW_INDEX_V1) and ZSTD or LZ4 on sparse families.
  5. Put all writes behind one helper that enforces the null rule and owns the dictionary.
  6. Audit every SingleColumnValueFilter for setFilterIfMissing(true).
  7. Switch to ROWCOL blooms only if the dominant read is a Get with explicit columns, then compare HFiles read per Get.
  8. Cap columns per row and use setBatch plus partial results for any row that can grow without bound.
Key takeaway: HBase stores only the cells you write, so missing attributes cost nothing, but every present cell repeats its row key, family, qualifier and timestamp, which can dwarf small values. Keep family names short, consider a qualifier dictionary, and enable block encoding and compression. Give null one meaning and clear attributes with deletes, set setFilterIfMissing on value filters, use ROWCOL blooms only for column-specific gets, and cap or batch rows that can grow without bound.