Most HBase tables are not loaded once. They are kept current: a change-data-capture stream from an operational database, a nightly feed of corrected records, a fleet of devices updating their status. Each change touches a few cells in a table of billions, and the job is to apply it correctly, cheaply and repeatably, and then to let downstream systems pick up only what changed. HBase makes the first part look easy because a write never needs a read, but the details that decide correctness are less visible: which timestamp each cell gets, what a delete really masks, what happens when the same event arrives twice or two events arrive out of order, and how a consumer can tell what is new.

This article works through those details. It compares three ways to model a change, shows why the cell timestamp is the most important decision you make, covers conditional updates with CheckAndMutate, delta rows with merge-on-read, and the three ways to read changes back out, then walks one order through a CDC pipeline. The client API itself is introduced in the HBase Java API, and counters under replay are covered in streaming ingest into HBase.

Three ways to model a change

There are three ways to represent a change, and a table usually mixes them by column family.

  • Upsert (last write wins). Each change writes the new value of the columns it touched. Reads return the latest version. This is the default and the cheapest: no read before write, and replaying the same change is harmless if the timestamp is deterministic.
  • Versioned history. The column family keeps several versions, so each change becomes a new version and the old values stay readable with readVersions(n) or an as-of time range. History is bounded by VERSIONS and TTL, and is removed by compaction, so it is a convenience, not an audit log; see TTL and versions.
  • Delta (merge on read). Each change writes a delta, such as +3 units of stock, as its own cell, and readers fold deltas onto a base value. Writes are idempotent if each delta is keyed by its event ID, and a periodic fold keeps reads cheap. Use it when the change is relative and replays must not double-apply.
ModelWrite costRead costSafe to replay?Use for
UpsertOne PutOne cellYes, with deterministic timestampsState: status, address, profile
VersionedOne PutOne cell; more for historyYes, same ruleRecent history, as-of reads
DeltaOne Put per eventBase plus pending deltasYes, keyed by event IDBalances, stock, aggregates
IncrementServer read-modify-writeOne cellNo; a replay double-countsApproximate counters only
Incremental changes in and out of HBaseSource databaseorders, accountsChange logCDC events, ordered per keyApplierevent -> Put / Delete / CAScommit time, seqHBase row order#9F2A-1042d:status PAID @t2 NEW @t1d:amount 4200 @t1m:seq 000000000017 @t2delete marker @t3 masks ts <= t3cell ts = source timeTime-range scancells in [from, to)Export toolstarttime / endtimeReplicationWAL shipping, continuousApplying changes IN depends on cell timestamps and conditions; reading changes OUT depends on the same timestamps.Choose the timestamp rule once, because it decides both.
Changes are applied with deterministic timestamps and read back out by time range, Export or replication.

Partial updates, NULLs and deletes

Because HBase stores every column as an independent cell, a Put that names two columns changes only those two; the rest of the row is untouched. That makes partial updates free: a CDC event that says only the status changed becomes a Put of d:status, with no read of the old row. The flip side is that HBase has no concept of NULL. A column set to null at the source must become a Delete of that column, not a Put of empty bytes, or readers will see an empty string where the source has no value.

Deletes need equal care, because there are three kinds and they mask differently. addColumn(f, q, ts) removes exactly one version. addColumns(f, q, ts) removes every version of the column at or below ts. A row Delete at ts removes every cell in the row at or below ts. The marker stays until a major compaction, and while it exists it hides matching cells even if they arrive after the delete. An applier that maps events to mutations looks like this:

// One CDC event -> one atomic row mutation. ts = source commit time (ms).
RowMutations toMutations(ChangeEvent e) throws IOException {
    byte[] row = rowKey(e.tableKey());
    long ts = e.commitTimeMillis();
    RowMutations rm = new RowMutations(row);
    if (e.op() == Op.DELETE) {
        rm.add(new Delete(row, ts));                  // masks everything <= ts
        return rm;
    }
    Put put = new Put(row, ts);
    Delete nulls = new Delete(row, ts);
    for (ColumnChange c : e.changedColumns()) {
        byte[] q = Bytes.toBytes(c.name());
        if (c.isNull()) nulls.addColumns(D, q, ts);   // NULL = remove the column
        else            put.addColumn(D, q, c.valueBytes());
    }
    put.addColumn(M, SEQ, Bytes.toBytes(String.format("%012d", e.sequence())));
    rm.add(put);
    if (!nulls.isEmpty()) rm.add(nulls);
    return rm;
}

The cell timestamp decides correctness

Every cell carries a timestamp, and by default the region server stamps it with its clock when the write arrives. For incremental updates that default is usually wrong, because arrival order is not change order: a retry, a reprocessed batch or two parallel appliers can deliver an older change after a newer one, and with server timestamps the older value wins simply by arriving last.

Setting the cell timestamp to the source's commit time fixes two problems at once. Replays become idempotent, because the same change writes the same cell coordinates, row, column and timestamp, again. And order stops mattering, because a read returns the version with the highest timestamp, regardless of when it arrived. The delete masking described above also works in your favour: a late update older than a delete stays hidden, which is what the source database did.

The rule brings three costs you must design for. TTL is computed from the cell timestamp, so a backfill of year-old changes into a family with a 90-day TTL is invisible the moment it lands. Re-creating a deleted row needs a newer timestamp than the delete; a source that re-inserts with an old clock value will see its insert masked until a major compaction removes the marker. Two changes in the same millisecond collide on identical coordinates, and which one survives is behaviour you should not rely on. Keep the source's own sequence number (a log position or transaction ID) in a separate column, as the applier above does with m:seq, and use it whenever millisecond order is not enough.

Conditional updates with CheckAndMutate

When a change is only valid if the row is in an expected state, use a conditional mutation. CheckAndMutate, the builder API introduced in HBase 2.4, applies a Put, Delete or RowMutations only if a condition on one cell holds, atomically within the row. The safest conditions are equality and absence, which avoids reasoning about byte-order comparisons: read the current sequence, decide, and write only if it has not changed in the meantime.

// Apply an event only if it is newer than what the row already holds.
boolean applyIfNewer(Table t, ChangeEvent e) throws IOException {
    byte[] row = rowKey(e.tableKey());
    byte[] incoming = Bytes.toBytes(String.format("%012d", e.sequence()));
    for (int attempt = 0; attempt < 5; attempt++) {
        Result r = t.get(new Get(row).addColumn(M, SEQ));
        byte[] current = r.getValue(M, SEQ);
        if (current != null && Bytes.compareTo(current, incoming) >= 0)
            return false;                                  // stale or duplicate: skip
        CheckAndMutate.Builder b = CheckAndMutate.newBuilder(row);
        CheckAndMutate cam = (current == null)
            ? b.ifNotExists(M, SEQ).build(toMutations(e))
            : b.ifEquals(M, SEQ, current).build(toMutations(e));
        if (t.checkAndMutate(cam).isSuccess()) return true;
        // someone else changed the row between our get and our write: retry
    }
    throw new IOException("contention on " + Bytes.toStringBinary(row));
}

The sequence is zero-padded text so that byte order matches numeric order. This costs a read per event and serialises writers on the row, so reserve it for entities where out-of-order application is genuinely harmful and timestamps cannot express the order, for example when the source has no reliable commit clock. For most CDC streams, partitioning the change log by row key, so one applier owns each key, plus source timestamps, is enough without any read.

Delta cells and merge on read

Relative changes such as "reserve 3 units" cannot be expressed as an upsert without first reading the current value, and Increment applies them twice on replay. The delta pattern writes each change as its own cell, with the event ID as the qualifier, in a dedicated family. Writing the same event twice rewrites the same cell. A fold job, or the reader, sums the deltas onto a base value, and the fold rewrites base and removes the folded deltas in one atomic row mutation:

# write path: idempotent, no read
put(row, family="x", qualifier=event_id, value=delta, ts=event_time)

# read path
base, through = get(row, "b:value"), get(row, "b:through")   # through = last folded ts
pending = [c for c in get_family(row, "x") if c.ts > through]
value = base + sum(c.value for c in pending)

# fold job, per row, run when the row has more than N pending deltas
cutoff = now() - SAFETY_LAG
folded = [c for c in pending if c.ts <= cutoff]
if folded:
    new_base = base + sum(c.value for c in folded)
    new_through = max(c.ts for c in folded)
    row_mutations(row,
        put("b:value", new_base), put("b:through", new_through),
        *[delete_column("x", c.qualifier, c.ts) for c in folded])
    # condition the mutation on b:through still equalling `through`

The safety lag protects against a late delta with an old event time arriving after the fold, which the reader would otherwise skip because it is older than b:through. Choose it longer than your worst expected delay, and route anything later to a correction path. Reads cost a little more than a single cell, and the fold keeps that bounded.

Reading changes back out

Downstream systems usually want only what changed since they last looked. HBase gives you three ways to get it.

  • Time-range scans. new Scan().setTimeRange(from, to) returns only cells whose timestamps fall in the half-open range. HFiles record the range of timestamps they contain, so files entirely outside the window can be skipped. Persist to as your watermark and start the next scan there.
  • The Export tool. hbase org.apache.hadoop.hbase.mapreduce.Export <table> <outdir> <versions> <starttime> <endtime> writes the cells in a time window to sequence files as a MapReduce job, suitable for a nightly incremental extract or for moving changes to another cluster.
  • Replication. For continuous propagation, cluster replication ships write-ahead-log edits to a peer as they happen; see HBase replication. Bulk-loaded HFiles bypass the write-ahead log, so they need replication of bulk loads enabled or a separate path; bulk load explains why.

Here the timestamp rule comes back. A time-range consumer is reading cell timestamps, and if those are source commit times, a change that commits at 10:00 but arrives at 10:20 lands below a watermark the consumer already advanced past at 10:10. Either keep the consumer's upper bound well behind the present (now minus the maximum ingest delay), or have the applier also write an ingest-time marker column with a server timestamp and drive extraction from that. Pick one and write it down, because the two halves of the pipeline must agree.

Worked example: one order through CDC

An orders table holds 50 million rows, mirrored from a relational database through a CDC stream of about 2 million changes a day. The row key is a hash prefix plus order ID, the log is partitioned by order ID, and the applier uses source commit time as the cell timestamp and the log position as m:seq. Follow one order:

  1. 10:00:01, insert. The applier writes status NEW, amount 4200 and seq 15, all at t1 = 10:00:01.000.
  2. 10:04:30, update. Status becomes PAID. The event names one column, so the applier writes d:status and seq 17 at t2; amount is untouched.
  3. A reprocessing replays the 10:00 insert at 10:06. It writes NEW at t1 again, to the same coordinates. Reads still return PAID, because t2 is newer. Nothing needed detecting.
  4. 10:30, delete. A row Delete at t3 hides every cell at or below t3.
  5. A delayed copy of the 10:04 update arrives at 10:35. It writes PAID at t2, which is below t3, so the delete masks it and the order stays deleted, matching the source.
  6. Downstream. An hourly job scans with a time range ending 30 minutes before now and exports the changed cells and delete markers it sees to the warehouse, advancing its watermark only after the export commits.

The table's column family has no TTL, since commit times can be old during a backfill, and keeps three versions so support staff can see recent status history. Deleted orders disappear physically at the next major compaction.

Failure modes

  • Old value wins. Server-assigned timestamps plus out-of-order delivery let a stale update overwrite a newer one. Use source timestamps or a sequence check.
  • Backfill vanishes. Source timestamps older than the family's TTL expire on arrival. Remove the TTL or route backfills to a family without one.
  • Row cannot be re-created. A delete marker newer than the re-insert hides it. Re-inserts must carry a later timestamp than the delete.
  • NULLs become empty strings. The applier wrote empty bytes instead of deleting the column.
  • Counters drift. Increments were replayed after a restart. Switch to delta cells keyed by event ID.
  • Consumers miss changes. The extraction watermark ran ahead of late-arriving cells. Lag the upper bound or extract by ingest time.
  • Hot rows under CheckAndMutate. Many writers contend on one row and retries pile up. Partition the change log by key so one writer owns each row.

What to do next

  1. Classify every column family as upsert, versioned or delta, and remove any Increment that must survive replay.
  2. Decide the cell timestamp rule (source commit time is the usual answer) and check every family's TTL against the oldest timestamp a backfill could carry.
  3. Write the event-to-mutation mapping, including NULLs and each delete type, and test it with duplicated and reordered events.
  4. Store the source sequence in its own column, and add CheckAndMutate only where order cannot be expressed by timestamps.
  5. Choose the extraction method and its watermark rule, and set the safety lag from measured ingest delay.
  6. Replay one day of changes twice into a test table and compare it with the source; the results must be identical.
Key takeaway: Incremental updates in HBase are correct when every change maps to a deterministic mutation. Use upserts for state, delta cells for relative changes, and source commit time as the cell timestamp so replays and reordering are harmless, but check TTLs, delete masking and extraction watermarks against that same rule, and reach for CheckAndMutate only where timestamps cannot express the order.