Moving an HBase table between clusters, or out to object storage and back, is one of the most common operational jobs on the platform, and the name of the tool that does it causes confusion. HBase ships two unrelated pairs. ExportSnapshot copies the immutable files that make up a snapshot from one filesystem root to another, without touching a RegionServer. The Export and Import MapReduce jobs scan a live table into sequence files of rows and replay them as writes. They differ in cost, consistency and flexibility, and picking the wrong one turns a two-hour copy into a two-day incident.
This article focuses on the file-level path: what ExportSnapshot does step by step (read from its source), how to size and throttle it, why checksum verification fails between HDFS and object stores, how the import side works, and how to verify the result. It closes with when the row-level tools are the better choice. Snapshot internals such as HFileLinks and the archive cleaner are covered in the snapshots article linked below.
The snapshot model in one minute
A snapshot is a manifest: the list of HFiles that made up each region of a table at one moment, plus the table descriptor. Taking one copies no data. Because HFiles are immutable, the manifest stays valid as long as the files are retained, and HBase keeps them in the archive directory after compaction while any snapshot still references them. That is what makes export cheap to reason about: copying a snapshot means copying a known, fixed set of files, while the source table continues to take writes.
It also defines what you get. The copy is consistent to the instant of the snapshot, per region (a flush snapshot flushes memstores first; edits after the snapshot are not included). If you need the target to catch up afterwards, pair the export with replication, as the migration article describes.
What ExportSnapshot does, step by step
Reading the tool's source makes its behaviour predictable. It first verifies the source snapshot's integrity and expiry. It then copies the snapshot directory (manifest and .snapshotinfo) from the source's .hbase-snapshot/<name> into the destination's working directory, .hbase-snapshot/.tmp/<name>. If you asked for a different target name with -target or for -reset-ttl, it rewrites .snapshotinfo there. Next it runs a map-only MapReduce job that copies every referenced HFile (and MOB file) into the destination's archive tree. Finally it renames the working directory to .hbase-snapshot/<name> and verifies the result.
The order is deliberate, and the source code says why: the references must exist before the files, or the destination's HFile cleaner would see unreferenced files in its archive and delete them. If the job fails, the driver deletes the .tmp directory and any final snapshot directory, but HFiles already copied stay in the destination archive. Because the mapper skips any file whose destination has the same length (and, with checksum verification on, the same checksum), a rerun to a backup root only copies what is missing. On a live cluster's root, that cluster's cleaner may remove the now-unreferenced files first. Only if the driver process itself dies does .tmp survive; the next run then refuses to start ("may be in-progress") until you remove it or pass -overwrite.
Options that matter
| Option | Effect | Notes |
|---|---|---|
-snapshot NAME | Snapshot to export | Required |
-copy-to URI | Destination HBase root | hdfs://..., s3a://..., any Hadoop filesystem |
-copy-from URI | Source root, default hbase.rootdir | Lets you pull from a remote root: the import side |
-target NAME | Name of the snapshot at the destination | Rewrites .snapshotinfo |
-mappers N | Number of map tasks | Default 1 + files / 10 |
-bandwidth MB | Throttle in MB/s | Per mapper; unthrottled by default |
-chuser / -chgroup / -chmod | Ownership and mode of copied files | chown needs HDFS superuser rights |
-overwrite | Delete an existing target snapshot or .tmp directory first | Archive files stay and are skipped |
-no-checksum-verify | Compare name and length only | Needed for most object stores |
-no-target-verify | Skip the final integrity check | Rarely a good idea |
-no-source-verify | Skip the initial source check | Present in 2.5 and later |
-reset-ttl | Do not carry the snapshot TTL to the target | Avoids an exported backup expiring |
Recent 2.6 and master builds also accept -custom-file-grouper and -file-location-resolver for locality-aware grouping. Run the tool with -help on your own build before relying on any option; the set has grown across releases.
Sizing and throttling, worked through
When -mappers is omitted, the job uses one mapper per ten files plus one (the snapshot.export.default.map.group setting). Files are grouped by size so each mapper gets a similar number of bytes, not a similar number of files. The -bandwidth limit is applied inside each mapper to its own input stream, so the aggregate rate is roughly mappers times bandwidth, and with no flag there is no limit at all.
Worked example. A 12 TB snapshot of an events table has 9,600 HFiles. Left alone, the tool would request 1 + 9,600 / 10 = 961 mappers, unthrottled, from a YARN cluster that shares disks and network with the RegionServers: an excellent way to raise read latency for every client. The link to the backup bucket is shared 10 Gbit/s, and the network team will allow about 400 MiB/s for this job. Choose 16 mappers with -bandwidth 25; the tool multiplies the value by 1,048,576, so that is 400 MiB/s or 419.4 MB/s. The copy then takes 12 x 10^12 bytes / 419,430,400 bytes/s = 28,610 seconds, just under 8 hours, plus the final verification. Fewer, bigger mappers also keep the request rate against an object store below the level at which it starts returning throttling errors.
Where to run it matters as much as how fast. Running on the source cluster's YARN reads HFiles locally but competes with serving traffic; running on the destination or a separate compute cluster moves the load away from production and reads remotely. For object-store imports, run on the destination.
Checksums across filesystems
By default every copied file is compared with its source by length and filesystem checksum. Between two HDFS clusters with the same block size this works. With different block sizes, HDFS's default per-block MD5-of-CRC checksums differ even for identical bytes; the tool's own error message suggests -Ddfs.checksum.combine.mode=COMPOSITE_CRC, which produces a block-size-independent file checksum. Object stores are a different case: an S3A path typically returns no HDFS-compatible checksum, and the 2.5, 2.6 and master code treats a missing or mismatched algorithm as a failed comparison, so the job fails with "Input and output filesystems are of different types".
The usual answer is -no-checksum-verify, which drops to name and length. Accept that consciously: it cannot detect a same-length corruption in transit. Compensate with the verification steps below, and keep the object store's own integrity features (such as checksums on upload) turned on.
Runbook: back up to S3, restore elsewhere
There is no separate import tool for snapshots. Importing is ExportSnapshot run with -copy-from pointing at the remote root and -copy-to at the local hbase.rootdir, followed by a clone or restore. The runbook below backs a table up to S3 and later brings it back on a different cluster.
# --- on the source cluster ---------------------------------------------
echo "snapshot 'prod:events', 'events-20261004'" | hbase shell -n
hbase org.apache.hadoop.hbase.snapshot.ExportSnapshot \
-snapshot events-20261004 \
-copy-to s3a://acme-hbase-backup/hbase \
-mappers 16 -bandwidth 25 \
-no-checksum-verify -reset-ttl
# inspect what landed (reads the manifest from the remote root)
hbase org.apache.hadoop.hbase.snapshot.SnapshotInfo \
-snapshot events-20261004 -remote-dir s3a://acme-hbase-backup/hbase -stats
# --- on the destination cluster, as a user allowed to chown -------------
hbase org.apache.hadoop.hbase.snapshot.ExportSnapshot \
-snapshot events-20261004 \
-copy-from s3a://acme-hbase-backup/hbase \
-copy-to hdfs://nn-b:8020/hbase \
-chuser hbase -chgroup hbase \
-mappers 16 -bandwidth 50 -no-checksum-verify
hbase shell -n <<'EOF'
list_snapshots 'events-.*'
create_namespace 'restore'
clone_snapshot 'events-20261004', 'restore:events'
EOFclone_snapshot creates a new table whose store files are links to the archived HFiles, so it is fast and copies nothing; compactions gradually rewrite the data into the table's own files. Use restore_snapshot only when you mean to roll an existing, disabled table back to the snapshot: it discards everything written since. The namespace must exist before cloning into it. If the copied files are owned by the user who ran the job rather than the HBase service user, compaction and cleanup on the new table can fail with permission errors; that is what -chuser and -chgroup are for.
Repeated exports and retention on a backup root
HFile names are unique and files are never modified, which gives repeated exports a useful property. Export tonight's snapshot to the same backup root as last night's, and every HFile that has not been compacted away since is already there with the same name and length, so the mapper skips it. Only files written or rewritten since the previous export are copied. For a table whose data mostly sits in large, rarely compacted files, a daily export can move a small fraction of the table. Watch the job's skipped-bytes counter to see the ratio.
The flip side is retention. A backup root on object storage has no HBase cluster attached, so nothing runs the archive cleaner there. Deleting an old snapshot's directory under .hbase-snapshot removes the manifest but leaves every HFile it referenced in archive, forever. Retention needs its own garbage collector: list the files referenced by the snapshots you keep (SnapshotInfo -files with -remote-dir prints them), and delete only archive files referenced by none. Run it while no export is writing to the root, or a file copied for an in-progress export can be collected before its manifest is renamed into place.
Verifying the copy
The tool's own target verification checks that every file the manifest references exists. It does not prove the bytes are right when checksums were skipped. Layer three cheap checks. Compare SnapshotInfo -stats output (file count and total size) on both sides. After cloning, run RowCounter on both tables, or on a key range for very large tables. For a stronger guarantee, HashTable on the source and SyncTable --dryrun against the clone compare hashed key ranges and report differences without changing anything. Record the results with the backup, so a restore drill can show what was verified and when.
Row-level Export and Import, and when to use them
The row-level pair works through the RegionServers instead of the filesystem. Export takes <tablename> <outputdir> [<versions> [<starttime> [<endtime>]]] and writes sequence files of results; Import <tablename> <inputdir> replays them as writes, or with -Dimport.bulk.output=/path produces HFiles for a bulk load instead. It also accepts a filter class through -Dimport.filter.class.
| Need | ExportSnapshot | Export / Import |
|---|---|---|
| Full table, point in time | Best fit: file copy, no RegionServer load | Slow; scans are not a single point in time across regions |
| A time window of edits | Not possible; whole snapshot only | Yes, via starttime and endtime |
| Filter or transform rows | No | Yes, with an import filter |
| Cost on the source | Disk and network reads only | Full scan load on RegionServers |
| Target differs in schema or split points | Clone recreates the source layout | Rows land in the target's layout |
In practice: use snapshots for backups, migrations and seeding replicas, and the row-level tools for slices, transformations, and moving data into a table whose design differs from the source.
Failure modes
- Unthrottled default. Hundreds of mappers with no bandwidth cap saturate shared links and disks.
- Leftover .tmp directory. A crashed driver blocks the next run until it is removed or
-overwriteis passed. - Checksum failures to object stores. Expected; use
-no-checksum-verifyand verify separately. - Expired exported snapshot. A source TTL carried to the backup can expire it; use
-reset-ttl. - Wrong file ownership. Files owned by the job user break compaction on the cloned table.
- Deleting the source snapshot too early. Until the export verifies, the snapshot is what keeps its HFiles in the archive.
- restore_snapshot by mistake. It overwrites a live table; clone to a new name and swap at the application layer.
Related reading
Start with HBase snapshots for manifests, HFileLinks and the archive cleaner, then cluster migration for export plus replication catch-up. See also disaster recovery, bulk loading and replication.
What to do next
- Run
ExportSnapshot -helpon your build and note which options it supports. - Export a small table to a scratch root with explicit mappers and bandwidth, and time it.
- Kill that MapReduce job halfway, rerun it to the same scratch root, and confirm the skipped-file counters rise.
- Import it on another cluster with -copy-from, clone to a new name, and verify with SnapshotInfo and RowCounter.
- Size the production export with the mappers times bandwidth arithmetic and agree the cap with the network owners.
- Schedule a quarterly restore drill that ends with HashTable and SyncTable --dryrun results attached to the ticket.