Most HBase disaster recovery plans fail in the same way: they choose one mechanism for the whole cluster. Either every table is replicated, and the WAN link saturates the first time a bulk job rewrites a large table, or everything relies on nightly snapshot exports, and the payments table quietly carries a day of exposure that nobody signed off on. A strategy is the step before the runbook. It decides, table by table and with arithmetic, which protection pattern each table gets, and checks that the chosen patterns fit the link, the standby cluster and the budget together.

This article is that selection step. It assumes you already know what RPO and RTO mean and how a standby cluster fails over; HBase Disaster Recovery, in depth covers the runbook, failback and drills, and HBase Backups covers the backup toolbox and retention. Here you measure two numbers per table, run a small planner, and read three results off it: the pattern per table, whether the WAN budget holds, and how long a backlog takes to drain after an outage. Every number in the worked example came from running the code shown.

The inputs a strategy needs

A DR pattern is chosen from four inputs per table, and two of them must be measured, not guessed.

  • Size on disk, after compression, from hdfs dfs -du -s -h /hbase/data/default/<table>. This drives how long a full copy takes and how much standby storage a pattern needs.
  • Peak WAL write rate in MB/s for the column families you would replicate. Replication ships WAL edits, so this is the bandwidth replication needs. Measure the peak hour, not the daily average; a batch job that rewrites a table at 02:00 is exactly when lag grows.
  • RPO, the data loss the owner will accept, in minutes.
  • RTO, the time allowed to be back in service, in minutes.

Per-table WAL rate is not reported directly, because a RegionServer's WAL interleaves edits from all its regions. Either multiply the table's peak-hour write request count by a sampled average mutation size, or replicate one table to a test peer and read the shipped-bytes metrics, which measures exactly what replication will ship.

Two inputs are global: the usable WAN bandwidth between sites, and the throughput your snapshot copies actually achieve. Never plan to fill the link; leave headroom for catch-up after an outage and for whatever else shares the circuit. The planner below uses 70 percent of a 1 Gbit/s link, about 87.5 MB/s.

Three patterns and what each costs

Three patterns cover almost every table. The planner chooses among them; the arithmetic for each is what makes the choice defensible.

PatternWorst-case data lossRecovery workSteady cost
Async replication to a warm standbyReplication lag, normally seconds; unbounded while the link is downRedirect clients; tables are already liveWAN bandwidth at peak WAL rate, a running standby
Snapshot export on a scheduleInterval plus the time to finish the copyRestore or clone the snapshot on the standby, then warm itPeriodic bulk transfer, standby storage per kept copy
Synchronous replicationClose to zero for acknowledged writesPromote the standbyWrite latency includes the remote round trip; version dependent

Async replication is the default for anything with an RPO under an hour. The source RegionServer tails its WALs and ships edits for families with REPLICATION_SCOPE => 1 to the peer; HBase replication architecture explains the path. The loss in a disaster is whatever was still queued, which is why the RPO is only as good as your lag alerting.

Snapshot export suits tables with RPOs measured in hours. A snapshot is a cheap metadata operation on the source (HBase Snapshots explains the file references), but ExportSnapshot then copies the referenced HFiles, skipping any the destination already holds with the same length. The worst case is a disaster just before the next snapshot completes copying: you lose one interval plus one copy time. That is why the planner requires the copy window to fit inside half the RPO.

Synchronous replication writes the remote WAL before acknowledging the client. It exists in some HBase 2.x releases, with peer states and a different failover procedure, but its availability and maturity differ by version and distribution. Treat an RPO under one minute as a design review, not a default.

A planner you can run

From measurements to a per-table DR patternMeasure per tablesize, peak WAL MB/sBusiness targetsRPO and RTO per tablePlannerrules + link budgetAsync replicationRPO = lag (seconds)Snapshot exportRPO = interval + copySync replicationRPO near zero, reviewEvery table also keeps local snapshotsreplication copies mistakes too: corruption needs a point in timeCheck: sum of replicated peak WAL rates must fit the WAN budget, with room to drain a backlog
The strategy is a function from measured numbers to a pattern per table, plus one global check on the shared WAN link. Local snapshots sit under every pattern because no replication method protects against a bad write.

The planner is deliberately small enough to read in a review. It encodes three rules: an RPO under a minute goes to a human, an RPO up to an hour gets replication, and anything slower gets snapshot export only if a full copy fits in half the RPO.

LINK_MB_S = 125.0      # 1 Gbit/s WAN, in MB/s
HEADROOM = 0.7         # plan to use at most 70% of the link
EXPORT_MB_S = 200.0    # measured ExportSnapshot throughput for the whole job

TABLES = [
    # name, size_gb, peak_wal_mb_s, rpo_min, rto_min
    ("orders",        900,  6.0,    5,   30),
    ("payments",      300,  2.5,    1,   15),
    ("sessions",      150,  9.0,  240,  240),
    ("events_raw",  18000, 22.0, 1440,  720),
    ("product_cat",    40,  0.2, 1440,  120),
]

def export_hours(size_gb):
    return size_gb * 1024 / EXPORT_MB_S / 3600

def choose(name, size_gb, wal, rpo, rto):
    if rpo < 1:
        return "sync-replication (review)"
    if rpo <= 60:
        return "async-replication"
    # snapshot shipping loses up to one interval plus one copy:
    # require the copy to fit in half the RPO
    if export_hours(size_gb) * 60 * 2 <= rpo:
        return "snapshot-export"
    return "async-replication"

plan = [(t, choose(*t)) for t in TABLES]
for t, pattern in plan:
    print(f"{t[0]:12} -> {pattern:18} export={export_hours(t[1]):.2f}h")

repl = sum(t[2] for t, pattern in plan if pattern == "async-replication")
budget = LINK_MB_S * HEADROOM
print(f"replicated peak WAL = {repl:.1f} MB/s; budget = {budget:.1f} MB/s")
print(f"if every table were replicated: {sum(t[2] for t in TABLES):.1f} MB/s")

outage_min = 30
backlog_mb = repl * outage_min * 60
spare = budget - repl
print(f"backlog after {outage_min} min = {backlog_mb / 1024:.1f} GB; "
      f"drains in {backlog_mb / spare / 60:.1f} min")

Running it prints:

orders       -> async-replication  export=1.28h
payments     -> async-replication  export=0.43h
sessions     -> snapshot-export    export=0.21h
events_raw   -> async-replication  export=25.60h
product_cat  -> snapshot-export    export=0.06h
replicated peak WAL = 30.5 MB/s; budget = 87.5 MB/s
if every table were replicated: 39.7 MB/s
backlog after 30 min = 53.6 GB; drains in 16.1 min

Worked example: reading the plan

Read the output line by line, because each line carries a decision you would otherwise make by instinct.

orders and payments go to replication on RPO alone. For payments, the one-minute target is legal only while lag stays under a minute, so the strategy must include an alert on the source metric ageOfLastShippedOp well below 60 seconds, and a stated position on what happens when the link is down. During an outage the real RPO is the outage length, and no pattern except synchronous replication changes that.

events_raw is the instructive one. Its owner accepts a day of loss, so instinct says nightly snapshot export. But the first copy of 18 TB at 200 MB/s takes 25.6 hours. Later exports skip HFiles the destination already has, yet every major compaction rewrites the files and makes the next copy full-size again, so the plan must assume the full window: longer than the interval, so the effective RPO is more than two days. Replicating it instead costs 22 MB/s of steady bandwidth, which is cheaper than a copy that never ends. The general lesson: large tables with a moderate write rate are often cheaper to replicate than to ship.

sessions is the opposite case. It writes at 9 MB/s, more than orders, but its RPO is four hours and it copies in 13 minutes, so snapshot export saves 9 MB/s of link budget.

The link check: 30.5 MB/s of replicated peak against an 87.5 MB/s budget fits, with 57 MB/s spare. Replicating everything would still fit at 39.7 MB/s, which tells you the link is not the binding constraint today. Keep that number: it is the first thing that changes when a new high-write table arrives.

The backlog check is the one most plans skip. After a 30-minute link outage, 53.6 GB is queued on the source, and it drains in 16.1 minutes using the spare capacity. While it drains, every replicated table's lag is far above its RPO. With 5 MB/s spare instead of 57, the same outage takes three hours to drain, while unshipped WALs pile up on the source HDFS: a common second incident after a network outage.

Turning the plan into configuration

Once the planner has assigned patterns, the configuration follows directly. The commands below are HBase 2.x shell syntax; adapt the ZooKeeper quorum and paths.

# Replicated tables: scope the families, then add one peer limited to those tables
alter 'orders',   {NAME => 'd', REPLICATION_SCOPE => '1'}
alter 'payments', {NAME => 'd', REPLICATION_SCOPE => '1'}
alter 'events_raw', {NAME => 'e', REPLICATION_SCOPE => '1'}
add_peer 'dr1', CLUSTER_KEY => "dr-zk1,dr-zk2,dr-zk3:2181:/hbase",
  TABLE_CFS => { "orders" => [], "payments" => [], "events_raw" => [] }
status 'replication'

# Snapshot-export tables: snapshot, then copy to the standby's HBase root
snapshot 'sessions', 'sessions-20261004-0600'
hbase org.apache.hadoop.hbase.snapshot.ExportSnapshot \
  -snapshot sessions-20261004-0600 \
  -copy-to hdfs://dr-nn:8020/hbase \
  -mappers 16 -bandwidth 100

# Periodic consistency check of a replicated table against the peer
hbase org.apache.hadoop.hbase.mapreduce.replication.VerifyReplication \
  --starttime=1791072000000 --endtime=1791075600000 dr1 orders

Two details matter. Create the tables on the standby with identical schemas before adding the peer, or edits will fail to apply. And measure ExportSnapshot throughput rather than trusting the -bandwidth flag: in the current source the cap, in MB/s, is applied inside each map task, so 16 mappers at 100 can move up to 1,600 MB/s in total. Check your release. The planner's EXPORT_MB_S should be what a real run achieved.

If your distribution ships the incremental hbase backup tool, it can replace scheduled full exports for large tables with WAL-based incrementals. It is not in every Apache release, so confirm it is present in yours before designing around it.

Standby capacity and RTO

Each pattern also costs standby capacity, and this is where strategies go over budget quietly.

  • Replication needs a standby that can absorb the peak write rate and, after failover, serve production reads. Size its RegionServers for the replicated tables' full read and write load, not a fraction, or failover becomes a slow overload. Storage equals the replicated tables' size.
  • Snapshot export needs storage for every retained copy. The first export is a full copy; later ones add the HFiles rewritten by compaction since the last, so budget one full copy plus that churn and measure the destination archive. Compute for the standby can be small until a disaster, which is the main saving of this pattern; the price is a longer RTO while you restore and warm the block cache.
  • Local snapshots on the primary are cheap at first and grow as compaction rewrites files: the snapshot keeps old HFiles alive in the archive directory. Set a retention and watch archive size.

RTO checks complete the plan. For a snapshot-export table, recovery is clone or restore, assignment and cache warm-up; time it on the standby with the real table size. If that time exceeds the RTO, either replicate the table or keep it pre-cloned on the standby. For replicated tables, RTO is mostly client redirection; read-heavy services that cannot wait even for that may want region replicas for single-server failures, which are a separate problem from site loss.

What the plan does not cover

The planner protects against site loss. A complete strategy also names the failures it does not cover, and how each is handled.

  • Logical corruption: a bad deploy that writes wrong values replicates in seconds. Only a point in time helps, so every table, replicated or not, needs local snapshots on a schedule that matches how quickly such bugs are noticed.
  • Schema drift: a column family added on the primary but not the standby makes replicated edits for it fail to apply. Make schema changes through a script that applies them to both clusters.
  • Dependencies: Kerberos principals, Ranger policies, coprocessor jars, Phoenix metadata and client configuration all need their own DR. A standby with data and no way to authenticate is not a standby.

Failure modes

The failure modes specific to strategy, as opposed to execution, are few and recurring.

  • One pattern for everything. Symptom: link saturation or an unowned day of exposure. Fix: per-table assignment.
  • Average instead of peak WAL rate. Symptom: lag that grows every night during batch loads. Fix: plan on the peak hour.
  • Export longer than the interval. Symptom: overlapping copies and an effective RPO of days. Fix: the half-RPO rule, or replication.
  • No drain headroom. Symptom: hours of lag and WAL buildup after a short network outage. Fix: keep spare link capacity and alert on WAL queue size.
  • Replication treated as backup. Symptom: a bad write destroys both copies. Fix: snapshots under every pattern.

What to do next

  1. Measure size and peak-hour WAL rate for every table, and write down an owner-approved RPO and RTO for each.
  2. Measure real ExportSnapshot throughput between your sites with a representative table.
  3. Run the planner with your numbers; challenge every table where the result surprises you.
  4. Check the link budget and the drain time after a 30-minute outage; buy headroom if drain exceeds your tightest RPO by a wide margin.
  5. Alert on replication lag below each replicated table's RPO, and on the source WAL queue size.
  6. Schedule local snapshots for every table, sized to how fast you detect bad writes.
  7. Time a restore on the standby for each snapshot-export table and compare it with the RTO.
  8. Then write the runbook and drills described in HBase Disaster Recovery, in depth, and re-run this plan monthly.
Key takeaway: An HBase DR strategy assigns a pattern to each table from measured numbers: async replication when the RPO is under an hour or a full copy cannot fit inside half the RPO, snapshot export when it can, and synchronous replication only after a design review. Check the sum of replicated peak WAL rates against the WAN budget, make sure spare capacity drains an outage backlog quickly, size the standby for the patterns chosen, and keep local snapshots under everything because replication copies bad writes too.