HBase gives you several ways to copy data: snapshots, ExportSnapshot, the Export and Import MapReduce jobs, CopyTable, cluster replication and, in newer releases, a backup utility. None of them is a backup program on its own. A program answers three questions before an incident: which failures must we survive, how far back and how precisely must we recover, and how do we know a restore will work.
This article builds that program. It maps failures to mechanisms, lays out a reference design combining hourly snapshots, daily off-cluster exports and continuously archived write-ahead logs, walks through a point-in-time recovery from a bad deploy, and covers verification, the metadata that lives outside your tables, retention sizing and restore drills. It assumes you know what a snapshot is; the internals are in the HBase snapshots deep dive, and the full-plus-incremental backup utility is covered in the HBase backup and restore architecture article.
Start from failures, not tools
Each mechanism survives some failures and not others. HDFS block replication survives a disk or node loss but copies corruption and deletes instantly. Region server crashes are handled by WAL splitting and replay, not by backups. Replication to a second cluster survives losing a site but faithfully replicates a dropped column family or a bug that overwrites good data, as the replication deep dive explains. Backups exist for the failures that propagate: human error, application bugs, malicious deletion and silent corruption.
| Failure | Survived by | Not survived by |
|---|---|---|
| Disk or node loss | HDFS replication | Nothing extra needed |
| Region server crash | WAL split and replay | Nothing extra needed |
| Site or cluster loss | Replication; exported snapshots off-site | On-cluster snapshots |
| Table dropped or truncated | Snapshots; exports | Replication (it replicates the damage) |
| Bad deploy writes wrong data for hours | Snapshot before the bug plus WAL replay up to it | Latest snapshot alone (it contains the bad data) |
| Compromised admin credentials | Immutable off-cluster copies in a separate account | Anything the attacker can delete |
| Silent corruption discovered weeks later | Long-retention exports with verification hashes | Short-TTL snapshots |
From that table come the two numbers every backup program needs per table: recovery point objective (RPO), how much recent data you may lose, and recovery time objective (RTO), how long a restore may take. A user-profile table might accept a 15-minute RPO and a 2-hour RTO; an audit table might require near-zero RPO but tolerate a day to restore.
The toolbox and what each piece costs
| Mechanism | What it captures | Cluster impact | Typical role |
|---|---|---|---|
| snapshot | References to current HFiles (optionally after a flush) | Seconds; storage grows as compaction replaces files | Fast rollback; source for exports |
| ExportSnapshot | A full copy of a snapshot's files to another filesystem or object store | MapReduce job; throttle with -bandwidth | Off-cluster daily copies |
| Archived WALs | Every edit in write order, with write time | Copy only rolled files; small | Closes the gap between snapshots |
| WALPlayer | Replays WAL edits for chosen tables and a time range | MapReduce job, or HFiles for bulk load | Point-in-time recovery |
| Export / Import | Cells in a time range, as sequence files | Full table scan | Logical copies across versions |
| Backup utility | Full images plus WAL-based incrementals, tracked in a system table | Managed by HBase | Integrated chains where your release ships it |
The backup utility shipped as a feature of the Apache HBase 3.0 line and exists as backports in some vendor 2.x distributions; if your release has it, it automates much of what follows. Everyone else builds the same capability from snapshots, exports and WALs, which is also the clearest way to understand what the utility does.
A reference design
The design uses three copies with three jobs. Hourly on-cluster snapshots with a short TTL give fast rollback for operator error. A daily ExportSnapshot of one of those snapshots to an off-cluster store, ideally object storage in a separate account with object locking, survives cluster loss and credential compromise. And a job copies each rolled WAL from the archive directory to the same off-cluster store every few minutes, so that edits since the last snapshot can be replayed.
The WAL copy is the part most teams miss. Once a WAL is rolled and no longer needed for recovery, HBase moves it to the oldWALs archive, where the log cleaner deletes it after a short time-to-live measured in minutes by default, unless a replication peer or the backup utility still holds it. If your copier falls behind the cleaner, your point-in-time window has a hole. Either raise the cleaner TTL to comfortably exceed the copier's worst lag, or monitor the copier's position against the oldest file in the archive.
#!/usr/bin/env bash
# daily off-cluster copy of one table, run from an edge node
set -euo pipefail
TABLE="shop:orders"
STAMP=$(date -u +%Y%m%dT%H%M)
SNAP="orders-daily-${STAMP}"
# 1. snapshot (flushes memstores first by default)
echo "snapshot '${TABLE}', '${SNAP}'" | hbase shell -n
# 2. copy it off-cluster, throttled to protect production
hbase org.apache.hadoop.hbase.snapshot.ExportSnapshot \
-snapshot "${SNAP}" \
-copy-to s3a://backup-acct-bucket/hbase \
-mappers 16 -bandwidth 200
# 3. record what we made, for the restore runbook
echo "${STAMP},${TABLE},${SNAP}" >> /var/lib/hbase-backup/catalog.csvRecent 2.x releases also let you set a TTL on a snapshot when creating it, so the master's cleaner removes expired snapshots automatically; check your version before relying on it, and keep the retention logic in your own catalog either way.
Worked incident: point-in-time recovery from a bad deploy
At 14:10 UTC a deploy starts writing a corrupted price field into the orders table. It is noticed at 16:40. The hourly snapshot from 14:00 is clean; every snapshot after it contains bad rows. The goal is the table as it stood at 14:09:59, without destroying the live table, which is still receiving good writes to other columns.
- Clone the 14:00 snapshot into a new table,
shop:orders_pitr, withclone_snapshot. Cloning is copy-on-write and fast; the live table is untouched. - Collect archived WALs covering 14:00 to 14:10 from the off-cluster archive, including a margin before 14:00, because the snapshot's flush did not happen at exactly 14:00:00.
- Replay those WALs into the clone, restricted to the orders table and to write times before the deploy, mapping the source table name to the clone.
- Verify the clone, then either swap applications to it or use it as the source to repair the affected cells in the live table.
# replay edits written between 13:55:00 and 14:09:59.99 UTC on 2026-09-29 into the clone;
# epoch milliseconds avoid any doubt about which time zone a date string is parsed in
hbase org.apache.hadoop.hbase.mapreduce.WALPlayer \
-Dwal.start.time=1790690100000 \
-Dwal.end.time=1790690999990 \
-Dmapreduce.map.speculative=false \
s3a://backup-acct-bucket/hbase-wal/2026-09-29/ \
shop:orders shop:orders_pitrReplaying the margin is safe because an HBase cell is identified by row, column and timestamp: re-applying an edit that the snapshot already contains writes the same cell again. Two caveats matter. The time filter uses the WAL entry's write time, not any timestamp the client set on the cell, so applications that set their own timestamps need care. And deletes replay too: a delete issued at 14:05 will be applied, which is correct for a point-in-time copy but surprising if you expected to recover that row. For large replays, WALPlayer can write HFiles through its bulk output option for one table at a time, which you then load as described in the bulk load deep dive.
Verify, or it is not a backup
Every copy needs a check that it is complete and correct. Three built-in tools cover most needs. RowCounter gives a fast sanity count. HashTable computes hashes over row ranges of a source table, and SyncTable compares a target against those hashes, reporting or fixing differences; with dry run on, it only reports. Comparing a restored clone against the live table for a stable key range, or against a clone of the same snapshot, gives strong evidence the restore path works.
# counts: quick sanity check on the restored clone
hbase org.apache.hadoop.hbase.mapreduce.RowCounter shop:orders_pitr
# hash a clone of the source snapshot, then compare the restored table against it
hbase org.apache.hadoop.hbase.mapreduce.HashTable --batchsize=32000 \
shop:orders_src_clone /backup-verify/hashes/orders
hbase org.apache.hadoop.hbase.mapreduce.SyncTable --dryrun=true \
/backup-verify/hashes/orders shop:orders_src_clone shop:orders_restoredRecord the counters from each run alongside the backup catalog entry. A backup with no verification record should be treated as unverified in any incident review.
The data that is not in your tables
A cluster rebuilt from table snapshots alone is missing things. Capture these separately, as text, on the same schedule as exports:
- Namespaces and table descriptors: column families, compression, block encoding, TTLs, versions and split points. Snapshots carry the table descriptor, but not every namespace setting.
- Access control: grants in the
hbase:acltable, or policies in an external system such as Ranger. - Quotas, region server groups and replication peer definitions.
- Application-layer metadata. Phoenix keeps its schema in
SYSTEM.CATALOGand maintains index tables; restoring a data table without its index tables, or at a different point in time, leaves indexes inconsistent until rebuilt. - The runbook itself, the catalog of backups and the credentials needed to read the off-cluster store, stored somewhere that survives the cluster.
Sizing retention and cost
A snapshot costs almost nothing when taken, because it references existing HFiles. It becomes expensive when compaction rewrites those files: the originals move to the archive and are kept while any snapshot references them. A table that fully compacts weekly and keeps hourly snapshots for a week can hold close to two full copies of its data. Keep on-cluster snapshots short-lived, 24 to 48 hours of hourlies is common, and move longer retention off-cluster.
For the off-cluster tier, a worked sizing: a 10 TB table, daily full exports retained 14 days, weekly exports retained 8 weeks, and WAL archives of about 200 GB per day retained 14 days. That is 14 plus 8, 22 full copies, or 220 TB, plus 2.8 TB of WALs. If that is too much, keep fewer dailies and rely on WAL replay from the most recent weekly for older points, trading storage for a longer restore. The WAL mechanics that make this possible are covered in the write-ahead log deep dive.
Restore drills and failure modes
A quarterly drill restores a real table from the off-cluster store into an isolated environment, replays WALs to a chosen time, verifies with the tools above and records the elapsed time against the RTO. Rotate who runs it, so the runbook is tested by someone who did not write it.
| Failure | Symptom | Prevention |
|---|---|---|
| WAL gap | Replay has missing minutes; recovered data is inconsistent | Cleaner TTL above copier lag; alert on copier position |
| Snapshot-only thinking | Every snapshot contains the bad data | Keep WALs to recover to a time before the bug |
| Restore overwrote production | restore_snapshot on the live table destroyed newer data | Always clone first; restore in place only by decision |
| Unthrottled export | Latency spikes during backup windows | Use -bandwidth and -mappers; schedule off-peak |
| Backups deletable by the attacker | Credentials that manage HBase can also delete copies | Separate account, write-once object locking |
| Stale indexes after restore | Phoenix queries return wrong rows | Restore index tables together or rebuild them |
What to do next
- List every table with its RPO, RTO and the failures in the failure table it must survive.
- Schedule hourly snapshots with a short retention and daily ExportSnapshot to a separate, write-once store.
- Build the WAL copier, set the log cleaner TTL above its worst lag, and alert when it falls behind.
- Export namespaces, descriptors, ACLs, quotas, peers and Phoenix catalog on the same schedule.
- Write the point-in-time runbook: clone, replay with WALPlayer time bounds, verify, then swap or repair.
- Add RowCounter and HashTable/SyncTable verification and record results in the backup catalog.
- Size retention with real compaction behaviour and move long retention off-cluster.
- Run a timed restore drill each quarter with a rotating operator and compare the result to the RTO.