HBase gives you several ways to copy data: snapshots, ExportSnapshot, the Export and Import MapReduce jobs, CopyTable, cluster replication and, in newer releases, a backup utility. None of them is a backup program on its own. A program answers three questions before an incident: which failures must we survive, how far back and how precisely must we recover, and how do we know a restore will work.

This article builds that program. It maps failures to mechanisms, lays out a reference design combining hourly snapshots, daily off-cluster exports and continuously archived write-ahead logs, walks through a point-in-time recovery from a bad deploy, and covers verification, the metadata that lives outside your tables, retention sizing and restore drills. It assumes you know what a snapshot is; the internals are in the HBase snapshots deep dive, and the full-plus-incremental backup utility is covered in the HBase backup and restore architecture article.

Advertisement

Start from failures, not tools

Each mechanism survives some failures and not others. HDFS block replication survives a disk or node loss but copies corruption and deletes instantly. Region server crashes are handled by WAL splitting and replay, not by backups. Replication to a second cluster survives losing a site but faithfully replicates a dropped column family or a bug that overwrites good data, as the replication deep dive explains. Backups exist for the failures that propagate: human error, application bugs, malicious deletion and silent corruption.

FailureSurvived byNot survived by
Disk or node lossHDFS replicationNothing extra needed
Region server crashWAL split and replayNothing extra needed
Site or cluster lossReplication; exported snapshots off-siteOn-cluster snapshots
Table dropped or truncatedSnapshots; exportsReplication (it replicates the damage)
Bad deploy writes wrong data for hoursSnapshot before the bug plus WAL replay up to itLatest snapshot alone (it contains the bad data)
Compromised admin credentialsImmutable off-cluster copies in a separate accountAnything the attacker can delete
Silent corruption discovered weeks laterLong-retention exports with verification hashesShort-TTL snapshots

From that table come the two numbers every backup program needs per table: recovery point objective (RPO), how much recent data you may lose, and recovery time objective (RTO), how long a restore may take. A user-profile table might accept a 15-minute RPO and a 2-hour RTO; an audit table might require near-zero RPO but tolerate a day to restore.

The toolbox and what each piece costs

MechanismWhat it capturesCluster impactTypical role
snapshotReferences to current HFiles (optionally after a flush)Seconds; storage grows as compaction replaces filesFast rollback; source for exports
ExportSnapshotA full copy of a snapshot's files to another filesystem or object storeMapReduce job; throttle with -bandwidthOff-cluster daily copies
Archived WALsEvery edit in write order, with write timeCopy only rolled files; smallCloses the gap between snapshots
WALPlayerReplays WAL edits for chosen tables and a time rangeMapReduce job, or HFiles for bulk loadPoint-in-time recovery
Export / ImportCells in a time range, as sequence filesFull table scanLogical copies across versions
Backup utilityFull images plus WAL-based incrementals, tracked in a system tableManaged by HBaseIntegrated chains where your release ships it

The backup utility shipped as a feature of the Apache HBase 3.0 line and exists as backports in some vendor 2.x distributions; if your release has it, it automates much of what follows. Everyone else builds the same capability from snapshots, exports and WALs, which is also the clearest way to understand what the utility does.

Advertisement

A reference design

The design uses three copies with three jobs. Hourly on-cluster snapshots with a short TTL give fast rollback for operator error. A daily ExportSnapshot of one of those snapshots to an off-cluster store, ideally object storage in a separate account with object locking, survives cluster loss and credential compromise. And a job copies each rolled WAL from the archive directory to the same off-cluster store every few minutes, so that edits since the last snapshot can be replayed.

The WAL copy is the part most teams miss. Once a WAL is rolled and no longer needed for recovery, HBase moves it to the oldWALs archive, where the log cleaner deletes it after a short time-to-live measured in minutes by default, unless a replication peer or the backup utility still holds it. If your copier falls behind the cleaner, your point-in-time window has a hole. Either raise the cleaner TTL to comfortably exceed the copier's worst lag, or monitor the copier's position against the oldest file in the archive.

An HBase backup program: three copies with three jobsproduction clustertables, hbase:meta, ACLshourly snapshotson-cluster, short TTLarchived WALscopied every few minutessnapshotcopy rolledoff-cluster storedaily ExportSnapshotWAL archivetime-stamped, immutableexportrestore environmentclone snapshot, replay WALs with WALPlayer, verify with HashTable and SyncTableconfig exportschemas, ACLs, quotas, peersSnapshots give fast rollback, exports survive cluster loss, archived WALs close the gap to a point in time.The restore environment proves the other three work.
Snapshots, exports and archived WALs each cover a different failure. The config export and the restore environment are part of the program, not afterthoughts.
#!/usr/bin/env bash
# daily off-cluster copy of one table, run from an edge node
set -euo pipefail
TABLE="shop:orders"
STAMP=$(date -u +%Y%m%dT%H%M)
SNAP="orders-daily-${STAMP}"

# 1. snapshot (flushes memstores first by default)
echo "snapshot '${TABLE}', '${SNAP}'" | hbase shell -n

# 2. copy it off-cluster, throttled to protect production
hbase org.apache.hadoop.hbase.snapshot.ExportSnapshot \
  -snapshot "${SNAP}" \
  -copy-to s3a://backup-acct-bucket/hbase \
  -mappers 16 -bandwidth 200

# 3. record what we made, for the restore runbook
echo "${STAMP},${TABLE},${SNAP}" >> /var/lib/hbase-backup/catalog.csv

Recent 2.x releases also let you set a TTL on a snapshot when creating it, so the master's cleaner removes expired snapshots automatically; check your version before relying on it, and keep the retention logic in your own catalog either way.

Worked incident: point-in-time recovery from a bad deploy

At 14:10 UTC a deploy starts writing a corrupted price field into the orders table. It is noticed at 16:40. The hourly snapshot from 14:00 is clean; every snapshot after it contains bad rows. The goal is the table as it stood at 14:09:59, without destroying the live table, which is still receiving good writes to other columns.

  1. Clone the 14:00 snapshot into a new table, shop:orders_pitr, with clone_snapshot. Cloning is copy-on-write and fast; the live table is untouched.
  2. Collect archived WALs covering 14:00 to 14:10 from the off-cluster archive, including a margin before 14:00, because the snapshot's flush did not happen at exactly 14:00:00.
  3. Replay those WALs into the clone, restricted to the orders table and to write times before the deploy, mapping the source table name to the clone.
  4. Verify the clone, then either swap applications to it or use it as the source to repair the affected cells in the live table.
# replay edits written between 13:55:00 and 14:09:59.99 UTC on 2026-09-29 into the clone;
# epoch milliseconds avoid any doubt about which time zone a date string is parsed in
hbase org.apache.hadoop.hbase.mapreduce.WALPlayer \
  -Dwal.start.time=1790690100000 \
  -Dwal.end.time=1790690999990 \
  -Dmapreduce.map.speculative=false \
  s3a://backup-acct-bucket/hbase-wal/2026-09-29/ \
  shop:orders shop:orders_pitr

Replaying the margin is safe because an HBase cell is identified by row, column and timestamp: re-applying an edit that the snapshot already contains writes the same cell again. Two caveats matter. The time filter uses the WAL entry's write time, not any timestamp the client set on the cell, so applications that set their own timestamps need care. And deletes replay too: a delete issued at 14:05 will be applied, which is correct for a point-in-time copy but surprising if you expected to recover that row. For large replays, WALPlayer can write HFiles through its bulk output option for one table at a time, which you then load as described in the bulk load deep dive.

Verify, or it is not a backup

Every copy needs a check that it is complete and correct. Three built-in tools cover most needs. RowCounter gives a fast sanity count. HashTable computes hashes over row ranges of a source table, and SyncTable compares a target against those hashes, reporting or fixing differences; with dry run on, it only reports. Comparing a restored clone against the live table for a stable key range, or against a clone of the same snapshot, gives strong evidence the restore path works.

# counts: quick sanity check on the restored clone
hbase org.apache.hadoop.hbase.mapreduce.RowCounter shop:orders_pitr

# hash a clone of the source snapshot, then compare the restored table against it
hbase org.apache.hadoop.hbase.mapreduce.HashTable --batchsize=32000 \
  shop:orders_src_clone /backup-verify/hashes/orders
hbase org.apache.hadoop.hbase.mapreduce.SyncTable --dryrun=true \
  /backup-verify/hashes/orders shop:orders_src_clone shop:orders_restored

Record the counters from each run alongside the backup catalog entry. A backup with no verification record should be treated as unverified in any incident review.

The data that is not in your tables

A cluster rebuilt from table snapshots alone is missing things. Capture these separately, as text, on the same schedule as exports:

  • Namespaces and table descriptors: column families, compression, block encoding, TTLs, versions and split points. Snapshots carry the table descriptor, but not every namespace setting.
  • Access control: grants in the hbase:acl table, or policies in an external system such as Ranger.
  • Quotas, region server groups and replication peer definitions.
  • Application-layer metadata. Phoenix keeps its schema in SYSTEM.CATALOG and maintains index tables; restoring a data table without its index tables, or at a different point in time, leaves indexes inconsistent until rebuilt.
  • The runbook itself, the catalog of backups and the credentials needed to read the off-cluster store, stored somewhere that survives the cluster.

Sizing retention and cost

A snapshot costs almost nothing when taken, because it references existing HFiles. It becomes expensive when compaction rewrites those files: the originals move to the archive and are kept while any snapshot references them. A table that fully compacts weekly and keeps hourly snapshots for a week can hold close to two full copies of its data. Keep on-cluster snapshots short-lived, 24 to 48 hours of hourlies is common, and move longer retention off-cluster.

For the off-cluster tier, a worked sizing: a 10 TB table, daily full exports retained 14 days, weekly exports retained 8 weeks, and WAL archives of about 200 GB per day retained 14 days. That is 14 plus 8, 22 full copies, or 220 TB, plus 2.8 TB of WALs. If that is too much, keep fewer dailies and rely on WAL replay from the most recent weekly for older points, trading storage for a longer restore. The WAL mechanics that make this possible are covered in the write-ahead log deep dive.

Restore drills and failure modes

A quarterly drill restores a real table from the off-cluster store into an isolated environment, replays WALs to a chosen time, verifies with the tools above and records the elapsed time against the RTO. Rotate who runs it, so the runbook is tested by someone who did not write it.

FailureSymptomPrevention
WAL gapReplay has missing minutes; recovered data is inconsistentCleaner TTL above copier lag; alert on copier position
Snapshot-only thinkingEvery snapshot contains the bad dataKeep WALs to recover to a time before the bug
Restore overwrote productionrestore_snapshot on the live table destroyed newer dataAlways clone first; restore in place only by decision
Unthrottled exportLatency spikes during backup windowsUse -bandwidth and -mappers; schedule off-peak
Backups deletable by the attackerCredentials that manage HBase can also delete copiesSeparate account, write-once object locking
Stale indexes after restorePhoenix queries return wrong rowsRestore index tables together or rebuild them

What to do next

  1. List every table with its RPO, RTO and the failures in the failure table it must survive.
  2. Schedule hourly snapshots with a short retention and daily ExportSnapshot to a separate, write-once store.
  3. Build the WAL copier, set the log cleaner TTL above its worst lag, and alert when it falls behind.
  4. Export namespaces, descriptors, ACLs, quotas, peers and Phoenix catalog on the same schedule.
  5. Write the point-in-time runbook: clone, replay with WALPlayer time bounds, verify, then swap or repair.
  6. Add RowCounter and HashTable/SyncTable verification and record results in the backup catalog.
  7. Size retention with real compaction behaviour and move long retention off-cluster.
  8. Run a timed restore drill each quarter with a rotating operator and compare the result to the RTO.
Key takeaway: An HBase backup program starts from the failures that propagate, human error, bad deploys, compromised credentials and slow corruption, which replication and HDFS copies do not survive. Combine short-lived on-cluster snapshots for rollback, daily exported snapshots in a separate write-once store for cluster loss, and continuously archived WALs so WALPlayer can replay to a point just before the damage. Capture the metadata that lives outside tables, verify every copy with RowCounter and HashTable with SyncTable, clone rather than restore in place, and prove the whole path with timed drills.