Relational modelling starts from the data and normalizes it so each fact lives once. HBase modelling starts from the queries and copies data until every important read is a single Get or one short Scan. That copying is denormalization, and in HBase it is not an optimization you add later: it is the design method, because the store has no joins, no built-in secondary indexes and no transactions across rows.

The price is that every copy must be kept correct by your code. This article explains the storage facts that make denormalization necessary, five patterns that cover most schemas, a worked order-history example with Java code, how to keep copies consistent without multi-row transactions, what the copies cost on disk, and the failure modes to plan for. It assumes you know what a rowkey and a column family are; HBase schema design covers those basics.

Why HBase forces denormalization

Four properties of HBase drive every pattern below.

  • Rows are sorted by rowkey and nothing else. A table is a sorted map from rowkey to cells, split into regions by key range. The only efficient lookups are by exact key and by key prefix or range. Any other predicate means a filtered scan that reads everything in range.
  • Atomicity stops at the row. A single Put or Delete on one row is atomic across all its column families, and Increment and Append update a cell atomically. The core client offers no transaction spanning two rows, let alone two tables.
  • Reads cost per row and per region. A Get visits one region server and checks the MemStore, block cache and possibly several HFiles. Reading 50 rows scattered across regions is 50 such operations, even when batched.
  • Writes are cheap. A write appends to the WAL and the MemStore; HFiles are written later by flushes and merged by compactions. Writing a fact three times costs three small sequential writes, which is usually far cheaper than the reads it saves.

So the trade is fixed: spend extra writes and storage to make each read a single row or a contiguous range, and accept that keeping the copies in agreement is your job.

Five denormalization patterns

Pattern 1: embed children as columns (wide row). Store an order and its line items in one row: d:status, d:total, then i:0001, i:0002 for items, each value a small serialized record. One Get returns the whole aggregate, and updating the order and an item together is atomic because it is one row. The limits are size and growth: a row is never split across regions, so an unbounded child set (all events for a device, all followers of a celebrity) produces huge rows that hurt compaction and memory, and every cell is bounded by hbase.client.keyvalue.maxsize on the client (10 MB by default). Use it when the children are bounded and read together.

Pattern 2: children as their own rows (tall table). Put each child in its own row with a composite key that starts with the parent: customerId|reverseTimestamp|orderId. A prefix scan on customerId| returns the customer's orders newest first, rows spread naturally across regions as they grow, and you can page with a start row. You lose the single-row atomicity between parent and children, so this suits large or unbounded collections that are read as ranges.

Pattern 3: duplicated index tables. To find orders by status or by SKU, write a second and third copy keyed by those attributes: status|reverseTimestamp|orderId in orders_by_status. The copy can hold just the primary key (a thin index that needs a second Get) or every field the query displays (a covering index that answers in one scan). Covering copies are faster and more expensive to keep correct.

Pattern 4: pre-aggregated counters. Instead of scanning orders to count them, Increment a counter row such as daily_counters: 2026-10-04|sku123 on every write. Reads become one Get. Increments are not idempotent, so a retried write double-counts; and a single global counter row is a hotspot, so bucket it by time or add a salt prefix and sum the buckets on read (see HBase hotspotting).

Pattern 5: copied reference data. Copy the product name and price into each order item at write time. That is not just a performance trick: the order should show the price paid, not today's price. Decide for each copied field whether it is a historical snapshot (never update) or a cached view of current data (must be refreshed), and write that decision into the schema document.

Worked example: an order history schema

Take an order service with four access patterns: show a customer's recent orders, open one order, list open orders by status for the operations team, and show daily units per SKU. Each pattern gets a table keyed for it, as in Figure 1.

QueryTableRowkeyRead
Customer's recent ordersorderscust|revTs|orderIdprefix scan, limit 20
One order with itemsorderssame row, families d and isingle Get
Open orders by statusorders_by_statusstatus|revTs|orderIdprefix scan
Daily units per SKUdaily_countersyyyy-mm-dd|skusingle Get

The reverse timestamp (Long.MAX_VALUE - epochMillis) makes newer orders sort first. The writer below applies one order event to all copies. It sets an explicit cell timestamp equal to the event's version, so replaying the same event writes the same cells, and a stale event cannot overwrite a newer one when reads use the latest version.

void apply(OrderEvent e) throws IOException {
    long ts = e.version();                      // monotonically increasing per order
    byte[] main = key(e.customerId(), rev(e.placedAt()), e.orderId());
    Put order = new Put(main, ts)
        .addColumn(D, STATUS, Bytes.toBytes(e.status()))
        .addColumn(D, TOTAL, Bytes.toBytes(e.totalCents()));
    for (Item it : e.items())
        order.addColumn(I, Bytes.toBytes(it.line()), it.serialize());
    orders.put(order);                          // atomic: one row

    // Index copies: write new entry, then remove the old one.
    byte[] idx = key(e.status(), rev(e.placedAt()), e.orderId());
    byStatus.put(new Put(idx, ts).addColumn(D, REF, main));
    if (e.previousStatus() != null)
        byStatus.delete(new Delete(key(e.previousStatus(), rev(e.placedAt()), e.orderId()), ts - 1));

    // Counters are not idempotent: only count each order once.
    if (e.isFirstVersion())
        for (Item it : e.items())
            counters.increment(new Increment(key(day(e.placedAt()), it.sku()))
                .addColumn(C, UNITS, it.quantity()));
}

Two choices in that code carry the design. The index delete uses timestamp ts - 1, so it removes only cells written at or before the previous version and cannot delete a newer index entry from a reordered event. And the counter update is guarded by isFirstVersion; a retry of the first event can still double-count, so counters are treated as approximate and corrected by the reconcile job, or the writer records processed event ids in the order row and checks them first.

One logical write, several physical rows: the denormalized write pathOrder serviceevent: order 9001 placedKafka logdurable, replayable1. appendIdempotent writerfixed timestamp = version2. consumeordersrow: cust|revTs|orderIdorders_by_statusrow: status|revTs|orderIdorders_by_skurow: sku|revTs|orderIddaily_countersIncrement per day bucket3a3b3c3dReconcile job (nightly)scan source, repair copiesNo transaction spans the four tables: each write is atomic only within its own row
Figure 1. A denormalized write path. The source of truth is the log; each HBase table is a derived copy keyed for one access pattern, written idempotently and repaired by a reconcile job.

Keeping copies consistent without transactions

Because the four writes are separate operations, a crash between them leaves copies disagreeing. Four techniques keep that manageable.

Write the source of truth first. Make one copy authoritative, usually the main row or, better, an upstream log such as Kafka. Derived copies are written after it, so every inconsistency is a missing or stale copy, never a copy pointing at data that does not exist.

Make every write idempotent. Explicit timestamps and full-row overwrites mean a replay converges to the same state. Feeding the writer from a log gives you replay for free: on any failure, re-consume from the last committed offset.

Verify on read. A thin index entry is a hint. When the status index says order 9001 is OPEN, the reader fetches the main row and drops the result if the main row now says SHIPPED, optionally deleting the stale index cell (read repair).

Reconcile in bulk. A scheduled job scans the authoritative table, recomputes every copy and counter for a key range, and repairs differences. Run it over recent ranges daily and the whole table less often. If you cannot run this job, you do not really know whether your copies are correct.

If you need index maintenance inside the database, Apache Phoenix maintains global secondary indexes for you, and coprocessors can trigger derived writes on the server side; both still write separate rows, so the same reasoning about partial failure applies (see Phoenix on HBase).

What copies cost on disk

Copies cost more than the values they hold, because HBase stores every cell with its full coordinates. A KeyValue carries a 4-byte key length, a 4-byte value length, a 2-byte row length, the row, a 1-byte family length, the family, the qualifier, an 8-byte timestamp and a 1-byte type: 20 fixed bytes plus row, family and qualifier. With a 32-byte rowkey, family d, an 8-byte qualifier and an 8-byte value, one cell is 20 + 32 + 1 + 8 + 8 = 69 bytes to store 8 bytes of data.

Multiply by copies. If an order has 12 cells in the main table and 4 in each of two covering indexes, that is 20 cells, about 1.4 KB before compression and before HDFS replication, against perhaps 200 bytes of actual values. Three habits contain it: keep family and qualifier names to one or two bytes, enable a data block encoding such as FAST_DIFF or PREFIX so repeated row prefixes are stored once per block, and compress HFiles. Short names matter less once encoding is on, but long rowkeys still cost in the block index and the cache. Group fields read together into the same family so a narrow read does not load wide ones (column family design).

Failure modes

  • Orphaned index entries. The index was written, the main row update failed. Readers that do not verify show phantom results. Verify on read and reconcile.
  • Missing index entries. The main row was written, the process died before the index. The order exists but cannot be found by status. Replay from the log or reconcile.
  • Reordered events. Two updates to one order arrive out of order and the older one wins. Use the event version as the cell timestamp, not the server clock.
  • Double-counted counters. Retries replay Increment. Deduplicate by event id or treat counters as approximate and recompute them.
  • Runaway wide rows. An embedded collection with no bound grows to gigabytes, a single region cannot split it, and compactions and reads slow down. Cap the collection or move it to a tall table.
  • Hot index prefixes. An index keyed by a low-cardinality value such as status sends all writes for OPEN to one region. Add a bucket prefix and scan all buckets.
  • Stale copied reference data. A field intended as a live view was copied as a snapshot. Document each copied field's intent and refresh live ones from change events.

Trade-offs

Every copy speeds up one read path and adds a write, storage and a correctness obligation. A useful rule is to denormalize for each query that must be fast and frequent, and to send rare analytical questions to a batch engine reading snapshots or exports instead of adding yet another index table. Covering copies beat thin copies when the read is latency-critical and fields change rarely; thin copies win when fields change often, since one main-row update is cheaper than rewriting every copy. Embedded children beat tall rows when the set is small and read whole; tall rows win as soon as the set can grow without bound. If most of your tables are index copies and you are writing your own index maintenance, measure whether Phoenix or a store with native secondary indexes would cost less engineering time.

What to do next

  1. List your top read queries with expected rate and latency target, and map each to exactly one table and rowkey prefix.
  2. For every copied field, record whether it is a historical snapshot or a live view, and how live views are refreshed.
  3. Route all writes through one idempotent writer fed by a replayable log, with explicit timestamps from event versions.
  4. Add verify-on-read to every thin index lookup and log the repairs you make.
  5. Build and schedule a reconcile job; report the number of repaired cells per run as a metric.
  6. Turn on data block encoding and compression, shorten family and qualifier names, and measure bytes per logical record before and after.
Key takeaway: HBase gives you sorted rowkeys and single-row atomicity, so you design one table or key prefix per important query and copy data into each. Treat one copy or an upstream log as the source of truth, write every copy idempotently with explicit timestamps, verify index hits on read, reconcile on a schedule, and keep cell names short and encoded because every copy repeats the full key.