Every Cassandra operator eventually learns that repair is not optional. Replicas drift apart whenever a write misses a node: a timeout, a dropped mutation, a hint that expired, a node that was down longer than the hint window. Read repair and hinted handoff close some of those gaps, but only anti-entropy repair guarantees that every replica of every range is compared and reconciled, and it has to do so before tombstones become eligible for purging, or deleted data comes back.
The mechanics of a repair session, Merkle trees, validation compaction, streaming and the repaired/unrepaired split, are covered in the repair architecture article. This article is about the other half: running repair as a program. What deadline you are working against, how to choose full or incremental per table, which scheduler to use now that Cassandra ships its own, how the Paxos step fits in, and how to know, from measurements rather than hope, that the cluster is actually converged.
Why repair is a deadline, not a chore
A delete in Cassandra writes a tombstone. Compaction may drop that tombstone only after gc_grace_seconds (864,000 seconds, ten days, by default) have passed. If one replica missed the delete and is not repaired within that window, the other replicas eventually purge their tombstones, and the next repair or read reconciliation sees the stale live value on the lagging replica as the newest data. The row reappears. This is the zombie-data failure, and it is silent: no error, no log line, just data a user deleted coming back.
So the rule that governs everything else is simple: every token range of every table must complete a successful repair at least once per gc_grace_seconds, with enough margin that one failed cycle can be re-run. The mechanics of tombstones and their purge rules explain why. Everything in a repair program, the scheduler, the parallelism, the choice of incremental or full, exists to meet that deadline without harming latency.
Two corollaries follow. A node that has been down longer than gc_grace_seconds must not simply rejoin; it may hold data that the rest of the cluster has already deleted and purged, so it should be wiped and replaced or rebuilt. And lowering gc_grace_seconds to shrink tombstone overhead is only safe if your measured repair cycle is comfortably shorter than the new value.
The 2026 toolbox
Four tools matter, and each has a job:
nodetool repairis the primitive. Since Cassandra 4.0 incremental repair is reliable (repair sessions mark SSTables as pending and anticompact at the start, not the end), and it is the default when--fullis not given. Useful flags:-prfor a node's primary range,-st/-etfor an explicit subrange,-osto reduce redundant streams, and-jfor job threads (at most 4).- Preview and validate.
--previewcomputes what a repair would stream without streaming it;--validatechecks that already-repaired data agrees across replicas. Both are measurement tools, and the cheapest way to answer whether the cluster is in sync. - Paxos repair. Lightweight transactions keep Paxos state per partition. Repair runs a Paxos step before data repair;
--paxos-onlyruns just that step and--skip-paxosskips it. Skipping is a deliberate exception, not a default: it exists for cases like a cluster that cannot complete Paxos repair during an incident. - A scheduler. Nobody runs repair by hand on a real cluster. Until recently the choice was Cassandra Reaper or home-grown cron. Now there is a third option inside the database itself.
# Is anything diverged right now? Estimate what a full repair would stream; repair nothing.
nodetool repair --preview --full orders
# Are the already-repaired SSTables still in agreement across replicas? (incremental users)
nodetool repair --validate orders
# One bounded session: the primary range of this node, full, fewer redundant streams.
nodetool repair --full -pr -os orders line_items
# One explicit subrange, the unit every scheduler actually works in.
nodetool repair --full -st -9223372036854775808 -et -9150000000000000000 orders line_items
# Clean up LWT (Paxos) state only, without touching table data -- e.g. before a topology change.
nodetool repair --paxos-only orders
Built-in Auto Repair (CEP-37)
CEP-37 added a native repair scheduler that runs inside every node, aware of topology and of its own repair history. According to the Cassandra documentation it was introduced for the 6.0 line and backported to 5.0.8. On 5.0.8 it must be switched on with the JVM property -Dcassandra.autorepair.enable=true before startup, and the documentation is explicit that this property is non-reversible: once enabled it cannot be disabled. Treat that as a one-way door and try it on a staging cluster first.
Configuration lives in an auto_repair block in cassandra.yaml, disabled by default, with global settings and per-type overrides for full, incremental and preview_repaired. The key knob is min_repair_interval (24 hours by default), the minimum time between repairs of the same data. Tables opt in or out through a table-level auto_repair property, and history is written to system_distributed.auto_repair_history. At runtime, nodetool getautorepairconfig, nodetool setautorepairconfig and nodetool autorepairstatus inspect and adjust it. Check the documentation for your exact version before relying on any other defaults; the splitter and parallelism settings have their own defaults that are worth reading rather than assuming.
# cassandra.yaml (6.0 line, or 5.0.8+ started with -Dcassandra.autorepair.enable=true)
auto_repair:
enabled: true # default is false
repair_type_overrides:
full:
enabled: true
min_repair_interval: 5d # comfortably inside a 10-day gc_grace
incremental:
enabled: false # turn on only after the migration steps below-- Opt a table out of incremental auto repair but keep full repair, and raise its priority.
ALTER TABLE orders.line_items
WITH auto_repair = {'incremental_enabled': 'false', 'full_enabled': 'true', 'priority': '1'};
-- Where the scheduler records what it has done, per node.
SELECT * FROM system_distributed.auto_repair_history LIMIT 20;
Choosing a scheduler
| Option | Strengths | Weaknesses | Choose it when |
|---|---|---|---|
| Auto Repair (built in) | No extra service; topology-aware; history in a system table; per-table opt-out | New; the 5.0.8 switch is one-way; fewer operators have run it at scale | You are on 6.0 or 5.0.8+ and can validate it in staging first |
| Reaper | Mature, UI and API, segment-level retries and pausing, widely run | Another service and backend to operate and upgrade | You already run it, or need its UI and multi-cluster view |
| cron + nodetool | Transparent, no dependencies | No coordination, no retries, easy to overlap sessions, silent failures | Tiny clusters or short-lived test environments only |
The deciding question is not features but ownership: who is paged when repair stops making progress? If the answer is nobody, no scheduler will save you. Pick one scheduler per cluster, never two; overlapping repairs from two tools on the same ranges multiply validation work and can fail each other's sessions.
Full or incremental, per table
Full repair builds Merkle trees over all data in the range every time. It is simple, has no persistent state beyond the SSTables, and its cost scales with data size. Incremental repair only compares data not yet marked repaired, so each run is cheaper, but it splits SSTables into repaired and unrepaired sets, depends on anticompaction, and needs occasional full repair or --validate to catch corruption in the repaired set.
| Table shape | Recommendation | Why |
|---|---|---|
| Large, append-mostly (events, time series with TWCS) | Incremental, plus periodic full or validate | Most data never changes; re-validating it every cycle is wasted I/O |
| Small or medium, heavily updated | Full | Cheap enough; avoids repaired/unrepaired SSTable split overhead |
| Delete-heavy with short gc_grace | Whichever finishes the cycle fastest, measured | The deadline dominates every other concern |
| LWT-heavy tables | Either, but never skip the Paxos step routinely | Paxos state must be repaired too |
Whatever you choose, use subrange sessions. A session that covers a large range builds coarse Merkle trees, where one differing partition marks a whole large leaf as different and causes over-streaming; small ranges keep trees fine-grained and failures cheap to retry.
Worked example: does the cycle fit?
Take a 24-node cluster in two datacenters of 12, replication factor 3 in each, about 4 TiB on disk per node, so 48 TiB per datacenter and roughly 16 TiB of distinct data per datacenter. Tables use the default ten-day gc_grace_seconds. Keep three days of margin so one failed cycle can be repeated, leaving a seven-day budget.
The throughput figure must come from your cluster, not a blog: time a few subrange sessions of known size from logs or scheduler history. Suppose they show about 150 GiB per hour per session and the latency SLO tolerates three concurrent sessions cluster-wide. The arithmetic below gives a full cycle of about 36 hours, well inside the 168-hour budget, so full repair every five days is affordable. If the answer were 150 hours, you would reach for incremental repair on the large append-mostly tables, more parallelism if latency allows, or a longer gc_grace_seconds, in that order.
# Does the repair cycle finish inside gc_grace_seconds with room to spare?
GC_GRACE_S = 864_000 # 10 days, the table default
MARGIN_S = 3 * 86_400 # keep 3 days for one failed cycle to be re-run
def cycle_ok(total_bytes_per_replica_set, sessions_in_parallel, bytes_per_hour_per_session):
hours = total_bytes_per_replica_set / (sessions_in_parallel * bytes_per_hour_per_session)
budget_h = (GC_GRACE_S - MARGIN_S) / 3600
return hours, budget_h, hours <= budget_h
# Worked example (measured numbers go here, not guesses):
TiB = 1024**4
hours, budget, ok = cycle_ok(
total_bytes_per_replica_set=48 * TiB / 3, # 48 TiB on disk per DC, RF 3 -> 16 TiB of distinct ranges
sessions_in_parallel=3,
bytes_per_hour_per_session=150 * 1024**3, # 150 GiB/h, taken from your own session logs
)
print(f"cycle {hours:.0f} h, budget {budget:.0f} h, ok={ok}") # cycle 36 h, budget 168 h, ok=TruePut the numbers on a dashboard. The single most useful repair metric is coverage age: for each table, the age of the oldest range since its last successful repair. Alert when it passes about 60 percent of gc_grace_seconds and page at 80 percent. That one alert catches a stalled scheduler, a failing table and a cluster that has outgrown its repair budget.
Moving an existing cluster to incremental repair
The Auto Repair documentation carries a warning worth repeating for any tool: enabling incremental repair on a cluster that has never run it can overwhelm nodes with anticompaction, because the first incremental run finds every SSTable unrepaired and splits them. A safe migration:
- Complete a full repair cycle of the table so replicas are known to agree.
- If your consistency requirements allow it, mark the existing SSTables as repaired with the offline
sstablerepairedsettool, node by node with the node stopped, so the first incremental run only sees new data. - Enable incremental repair for one table, watch pending compactions, disk headroom and read latency for a full cycle, then widen.
- Keep a periodic full repair or
--validaterun on a longer interval; incremental repair assumes the repaired set stays correct, and only a full comparison proves it.
Failure modes
- Stalled scheduler, green dashboards. Repair silently stops (expired credentials, a stuck lock, a node that fails every validation) and nothing alerts until zombie data appears. Coverage age is the defence.
- Disk exhaustion from streaming. Over-streaming on coarse ranges and anticompaction both need free space. Keep headroom, use small subranges, and watch pending compactions per node during repair windows.
- Repair overlapping topology changes. Bootstrap, decommission and replacement change range ownership mid-session. Pause repair during topology changes and resume afterwards; streaming during bootstrap already competes for the same bandwidth.
- Latency spikes from validation. Validation compaction reads every SSTable in the range. Cap concurrency, keep
-jlow, and schedule the heaviest tables away from peak. - Hints masking a lagging node. Hints cover outages shorter than the hint window; hinted handoff is not a substitute for repair and gives no guarantee after the window passes.
- Rejoining a stale node. A node down longer than gc_grace_seconds rejoining with its old data is the fastest route to resurrected rows. Replace it instead.
Trade-offs in one place
More repair parallelism finishes cycles sooner but costs tail latency. Incremental repair saves I/O but adds state that can itself go wrong. A longer gc_grace_seconds gives repair more slack but keeps tombstones longer, which slows reads over delete-heavy partitions. The built-in scheduler removes a moving part but, on 5.0.8, cannot be turned back off. Each trade-off should be settled with a measurement: cycle time, coverage age, p99 latency during repair and tombstone counts per read. The broader operations model describes how repair fits alongside backups, upgrades and capacity work.
What to do next
- List every table with its gc_grace_seconds and last successful repair time; that is your coverage-age baseline.
- Time a handful of subrange sessions and compute your full-cycle hours with the arithmetic above.
- Add coverage-age alerts at 60 and 80 percent of gc_grace_seconds.
- Pick exactly one scheduler; if you are considering Auto Repair on 5.0.8, trial it in staging, remembering the enable flag is one-way.
- Classify tables as full or incremental; migrate incremental candidates one at a time, full cycle first.
- Run
nodetool repair --previewweekly on a sample keyspace and trend the estimated streaming volume. - Write the runbook entry: pause repair for topology changes, and replace rather than rejoin nodes down longer than gc_grace_seconds.