Replication is not a backup. A Cassandra cluster with replication factor 3 survives a dead disk or node without anyone noticing, but it replicates a DROP TABLE, a buggy deploy that overwrites a million rows, or a mass delete to every replica within milliseconds. Backups exist for those cases, and for losing a whole cluster or region.
Cassandra makes backups cheap, because its data files are immutable. It also makes restores subtle, because every node owns different token ranges and holds its own copy of history. This article explains what a snapshot really is on disk, how incremental backups and commit log archiving extend it, how to ship files off the node, three restore paths and the correctness traps that make a restore quietly wrong. Command options are taken from the Cassandra 4.1 reference.
What you are protecting against
| Event | Does replication help? | What recovers it |
|---|---|---|
| Disk or node failure | Yes | Replace the node; it streams from replicas |
| DROP or TRUNCATE by mistake | No, replicated everywhere | Local auto snapshot, then import |
| Bad application writes | No | Point-in-time restore to a side cluster, then copy good rows back |
| SSTable corruption on one node | Mostly | Scrub or delete the file, then repair from replicas |
| Cluster or region lost | Only with a second data centre | Off-site snapshots restored onto new nodes |
| Ransomware or credential compromise | No | Off-site copies the cluster's credentials cannot delete |
Two numbers decide the design: the recovery point objective (how much recent data you can lose) and the recovery time objective (how long a restore may take). Nightly snapshots give an RPO of up to a day. Incremental backups shrink it to the flush interval. Commit log archiving can get close to zero, at real operational cost.
How a snapshot works on disk
Cassandra writes go to the commit log and a memtable, and memtables are flushed to SSTables that are never modified afterwards; compaction writes new files and deletes old ones. The file layout is covered in SSTable format architecture.
Because files never change, a snapshot does not copy data. nodetool snapshot first flushes memtables, then creates a hard link to every live SSTable component in data/<keyspace>/<table>-<table_id>/snapshots/<tag>/. A hard link is a second directory entry for the same inode, so it is instant and costs no space at first. Each snapshot directory also gets a schema.cql with the table definition and a manifest.json listing the files.
The cost appears later. When compaction replaces the original SSTables, the snapshot's links become the only owners of the old bytes, so disk use grows with the rate of compaction. A week of daily snapshots on a table that fully recompacts can hold several times the live data size. nodetool listsnapshots reports each snapshot's true size, which is the space you would free by clearing it.
The snapshot commands
nodetool snapshot -t pre-migration app_ks # one keyspace, tag "pre-migration"
nodetool snapshot -t orders-only -kt app_ks.orders # specific keyspace.table pairs
nodetool snapshot -t hourly --ttl 2d app_ks # expires automatically (4.1+)
nodetool snapshot -sf -t noflush app_ks # skip the flush: misses memtable data
nodetool listsnapshots # tag, table, true size, size on disk
nodetool clearsnapshot -t pre-migration -- app_ks # delete one tagSnapshots are per node and are taken whenever that node gets the command. Running it on twelve nodes from a loop produces twelve snapshots a few seconds apart, not one consistent cut of the cluster. That is acceptable because replicas are reconciled afterwards by repair, but it means a restore is never a precise instant in time.
Cassandra also snapshots for you. With auto_snapshot enabled in cassandra.yaml, which is the default, TRUNCATE and DROP TABLE take a snapshot first, so the most common operator error is recoverable locally. Those snapshots are kept until you clear them, unless auto_snapshot_ttl is set (4.1 and later, for example 30d), so a cluster that drops and recreates tables in tests slowly fills its disks; monitor snapshot true size.
Incremental backups
Setting incremental_backups: true in cassandra.yaml, or running nodetool enablebackup, makes Cassandra hard-link every newly flushed SSTable into the table's backups/ directory. nodetool statusbackup shows whether it is on.
Two facts shape how you use it. First, only flushed SSTables are linked, not the outputs of compaction, so incremental files are deltas against a snapshot: a restore needs the most recent full snapshot plus every incremental file taken after it. Second, Cassandra never deletes anything in backups/. Your shipping job must upload the files and then remove them, or the directory grows without bound and pins data that compaction already removed.
A typical schedule is a weekly snapshot, incremental files uploaded every few minutes, and a restore that replays the snapshot plus incrementals. The files carry no schema, so keep the snapshot's schema.cql with them.
Shipping backups off the node
A snapshot on the same disk does not survive the loss of that disk. Every production scheme copies files to object storage or another site and records enough identity to put them back. For each node you need the files, the node's tokens (with vnodes, a list of them), the schema and the Cassandra version. The script below is a minimal nightly job:
#!/usr/bin/env bash
# Nightly per-node backup: snapshot, record identity, upload, then drop the local snapshot.
set -euo pipefail
TAG="nightly-$(date -u +%Y%m%dT%H%M%SZ)"
HOST_ID=$(nodetool info | awk -F': ' '/^ID/ {print $2}')
DEST="s3://acme-cass-backups/prod/${HOST_ID}/${TAG}"
DATA=/var/lib/cassandra/data
nodetool snapshot -t "$TAG" app_ks # flushes, then hard-links
nodetool info -T | grep '^Token' | awk '{print $3}' > /tmp/tokens.txt
cqlsh -e "DESCRIBE KEYSPACE app_ks" > /tmp/schema.cql
# Upload only this tag's directories, keeping keyspace/table-id in the key.
cd "$DATA"
for dir in app_ks/*/snapshots/"$TAG"; do
aws s3 cp --recursive "$dir" "${DEST}/${dir%/snapshots/*}/"
done
aws s3 cp /tmp/tokens.txt "${DEST}/tokens.txt"
aws s3 cp /tmp/schema.cql "${DEST}/schema.cql"
nodetool clearsnapshot -t "$TAG" -- app_ks # free the hard links locallyKeying objects by host ID and keeping tokens next to the data is what makes a same-topology restore possible later. Throttle uploads so they do not compete with client traffic, write to a bucket the cluster's credentials cannot delete from, and encrypt at rest.
Most teams do not maintain this script themselves. Cassandra Medusa, an open-source tool from The Last Pickle and Spotify, wraps the same steps. medusa backup snapshots, uploads and clears one node; backup-cluster runs it across the cluster; restore-node and restore-cluster put data back; list-backups, verify and purge manage the catalogue. Its default mode is differential: because SSTables are immutable, a file already uploaded is referenced rather than copied again, so each backup uploads only new files. It supports S3 and S3-compatible stores, Google Cloud Storage, Azure Blob Storage and local storage.
Point-in-time recovery with commit log archiving
Snapshots and incremental files only contain flushed data. To recover to a specific moment, for example one minute before a bad deploy, Cassandra can archive every commit log segment and replay archived segments on restart. The mechanics of segments and replay are in commit log architecture. It is configured per node in conf/commitlog_archiving.properties:
# conf/commitlog_archiving.properties (every node)
archive_command=/bin/cp %path /backup/commitlog/%name
restore_command=/bin/cp -f %from %to
restore_directories=/restore/commitlog
restore_point_in_time=2026:09:30 14:05:00archive_command runs for each segment, with %path and %name substituted. On restore, segments from restore_directories are copied into place with restore_command and replayed, applying only mutations with a timestamp at or before restore_point_in_time, in the format yyyy:MM:dd HH:mm:ss (milliseconds and microseconds forms also exist).
The cutoff uses write timestamps, which clients can set themselves, so a client that writes with future timestamps defeats it. You also need every node's segments since that node's base snapshot, which is a lot of files to keep ordered. Treat point-in-time recovery as a practised procedure, restored into a side cluster, rather than a switch you flip in an emergency.
Three restore paths
| Situation | Method | Why |
|---|---|---|
| Same cluster, same table, lost data | Put files back and run nodetool import | Tokens unchanged; each node reloads its own files |
| New cluster, identical topology | Same initial_token values per node, recreate schema, import | Each new node owns exactly the ranges in its files |
| Different node count or tokens | sstableloader | Streams every row to whichever nodes own it now |
nodetool import (4.0 and later) loads SSTables from a directory into a live table. It verifies the files and, by default, checks that the tokens in them belong to this node. Useful options are -cd to copy rather than move the files, -t to skip the token check and -e for an extended verify. It replaces the older nodetool refresh.
Worked example: at 14:02 someone runs TRUNCATE app_ks.orders in production. Because auto_snapshot is on, every node has a snapshot whose name starts with truncated-. On each node:
# Recover app_ks.orders after an accidental TRUNCATE (auto_snapshot is on by default).
# Run on EVERY node, because every node took its own snapshot.
nodetool listsnapshots | grep orders # find the truncated-... tag
SNAP=$(ls -d /var/lib/cassandra/data/app_ks/orders-*/snapshots/truncated-* | tail -1)
# -cd copies the files instead of moving them, so the snapshot stays intact
nodetool import -cd app_ks orders "$SNAP"
# after all nodes are done, run on EVERY node (-pr covers only its primary ranges)
nodetool repair -pr app_ks ordersThe repair at the end matters, because each node's snapshot was taken at a slightly different moment. Writes after 14:02 are still in the table and merge with the restored rows by timestamp. Repair internals are in repair architecture.
For a new cluster with the same topology, set each new node's initial_token in cassandra.yaml to the saved token list of one old node, start the cluster empty, apply the schema, then import that old node's files on its twin. For a different topology use sstableloader -d host1,host2 /restore/app_ks/orders, where the last two path components must be the keyspace and table. The loader reads each file and streams rows to their current owners, so it works for any node count at the cost of network traffic and time. Streaming behaviour is described in bootstrap and streaming.
Traps that make a restore quietly wrong
- Resurrected deletes. A delete is a tombstone that compaction purges after
gc_grace_seconds(default 10 days). Restore a backup older than that into a live cluster and the deleted rows come back, because the tombstone that would have hidden them is gone. Restore old backups into a side cluster and copy selected rows instead. - Schema drift. Files written before a column was dropped or its type changed may not load cleanly. Restore with the schema saved alongside the backup, then migrate.
- Version mismatch. SSTable formats change between major versions. Restore onto the version that wrote the files, or check that the target can read them.
- Skipped flush.
-sfsnapshots miss whatever was only in the memtable. - Moved, not copied. Without
-cd, import moves files out of the snapshot directory, so a failed import can consume your only local copy. - Silent disk growth. Forgotten snapshots and an unshipped
backups/directory pin data that compaction removed. Alert on snapshot true size.
Compaction strategy affects all of this, because it decides how often files are rewritten and therefore how fast snapshots diverge from live data; see compaction strategies.
Operational guidance
- Write down RPO and RTO per keyspace; not every keyspace needs commit log archiving.
- Schedule snapshots outside peak hours and throttle uploads.
- Keep backups in a separate account or with object lock, so compromised cluster credentials cannot delete them.
- Store tokens, schema and version with every node's files.
- Clear local snapshots after upload, and alert when snapshot true size passes a set share of disk.
- Restore into a scratch cluster on a schedule, time it, and compare row counts or checksums on sample partitions. A backup that has never been restored is a hypothesis.
Choosing a scheme
| Scheme | Typical RPO | Storage cost | Operational burden |
|---|---|---|---|
| Nightly snapshots only | Up to 24 hours | Low; full files each time unless deduplicated | Low: one cron job and a restore runbook |
| Snapshots plus incremental backups | Minutes (flush interval) | Medium; many small files | Medium: the backups/ directory must be shipped and cleared |
| Plus commit log archiving | Seconds to minutes | High; every segment kept | High: ordered segment sets per node and a practised replay |
| Medusa, differential | As scheduled | Low; unchanged SSTables are not re-uploaded | Low to medium: one more tool to run and upgrade |
Most clusters are well served by Medusa or a snapshot job, with incremental backups for keyspaces whose data is expensive to recreate. Reserve commit log archiving for data where losing minutes has a real cost and you can staff the rehearsals.
What to do next
- Confirm auto_snapshot is on and alert on snapshot true size.
- Ship snapshots off-node with tokens, schema and version, or adopt Medusa.
- Decide per keyspace whether incremental backups or commit log archiving are worth their cost.
- Write the three restore runbooks from this page.
- Do a timed restore into a scratch cluster this month and every quarter after.