Every Cassandra operator eventually learns that repair is not optional. Replicas drift apart whenever a write misses a node: a timeout, a dropped mutation, a hint that expired, a node that was down longer than the hint window. Read repair and hinted handoff close some of those gaps, but only anti-entropy repair guarantees that every replica of every range is compared and reconciled, and it has to do so before tombstones become eligible for purging, or deleted data comes back.

The mechanics of a repair session, Merkle trees, validation compaction, streaming and the repaired/unrepaired split, are covered in the repair architecture article. This article is about the other half: running repair as a program. What deadline you are working against, how to choose full or incremental per table, which scheduler to use now that Cassandra ships its own, how the Paxos step fits in, and how to know, from measurements rather than hope, that the cluster is actually converged.

Advertisement

Why repair is a deadline, not a chore

A delete in Cassandra writes a tombstone. Compaction may drop that tombstone only after gc_grace_seconds (864,000 seconds, ten days, by default) have passed. If one replica missed the delete and is not repaired within that window, the other replicas eventually purge their tombstones, and the next repair or read reconciliation sees the stale live value on the lagging replica as the newest data. The row reappears. This is the zombie-data failure, and it is silent: no error, no log line, just data a user deleted coming back.

So the rule that governs everything else is simple: every token range of every table must complete a successful repair at least once per gc_grace_seconds, with enough margin that one failed cycle can be re-run. The mechanics of tombstones and their purge rules explain why. Everything in a repair program, the scheduler, the parallelism, the choice of incremental or full, exists to meet that deadline without harming latency.

Two corollaries follow. A node that has been down longer than gc_grace_seconds must not simply rejoin; it may hold data that the rest of the cluster has already deleted and purged, so it should be wiped and replaced or rebuilt. And lowering gc_grace_seconds to shrink tombstone overhead is only safe if your measured repair cycle is comfortably shorter than the new value.

A repair program: one deadline, one scheduler, three feedback signalsgc_grace_secondsthe hard deadline per tablePolicyfull / incremental per tableSchedulerAuto Repair, Reaper or cronSubrange sessionssmall token rangesPaxos repair stepLWT state firstData repairvalidate, then streamCoverage ageoldest unrepaired rangePreview / validatedivergence trendCost signalspending compactions, diskdeadlinepolicyalertsThe deadline and the policy feed the scheduler; the scheduler emits small sessions;coverage, divergence and cost measurements feed back into how aggressively it runs.
A repair program: gc_grace_seconds sets the deadline, a per-table policy picks the repair type, a scheduler turns both into small sessions, and three measurements (coverage age, divergence, cost) feed back into how hard it runs.

The 2026 toolbox

Four tools matter, and each has a job:

  • nodetool repair is the primitive. Since Cassandra 4.0 incremental repair is reliable (repair sessions mark SSTables as pending and anticompact at the start, not the end), and it is the default when --full is not given. Useful flags: -pr for a node's primary range, -st/-et for an explicit subrange, -os to reduce redundant streams, and -j for job threads (at most 4).
  • Preview and validate. --preview computes what a repair would stream without streaming it; --validate checks that already-repaired data agrees across replicas. Both are measurement tools, and the cheapest way to answer whether the cluster is in sync.
  • Paxos repair. Lightweight transactions keep Paxos state per partition. Repair runs a Paxos step before data repair; --paxos-only runs just that step and --skip-paxos skips it. Skipping is a deliberate exception, not a default: it exists for cases like a cluster that cannot complete Paxos repair during an incident.
  • A scheduler. Nobody runs repair by hand on a real cluster. Until recently the choice was Cassandra Reaper or home-grown cron. Now there is a third option inside the database itself.
# Is anything diverged right now? Estimate what a full repair would stream; repair nothing.
nodetool repair --preview --full orders

# Are the already-repaired SSTables still in agreement across replicas? (incremental users)
nodetool repair --validate orders

# One bounded session: the primary range of this node, full, fewer redundant streams.
nodetool repair --full -pr -os orders line_items

# One explicit subrange, the unit every scheduler actually works in.
nodetool repair --full -st -9223372036854775808 -et -9150000000000000000 orders line_items

# Clean up LWT (Paxos) state only, without touching table data -- e.g. before a topology change.
nodetool repair --paxos-only orders
Advertisement

Built-in Auto Repair (CEP-37)

CEP-37 added a native repair scheduler that runs inside every node, aware of topology and of its own repair history. According to the Cassandra documentation it was introduced for the 6.0 line and backported to 5.0.8. On 5.0.8 it must be switched on with the JVM property -Dcassandra.autorepair.enable=true before startup, and the documentation is explicit that this property is non-reversible: once enabled it cannot be disabled. Treat that as a one-way door and try it on a staging cluster first.

Configuration lives in an auto_repair block in cassandra.yaml, disabled by default, with global settings and per-type overrides for full, incremental and preview_repaired. The key knob is min_repair_interval (24 hours by default), the minimum time between repairs of the same data. Tables opt in or out through a table-level auto_repair property, and history is written to system_distributed.auto_repair_history. At runtime, nodetool getautorepairconfig, nodetool setautorepairconfig and nodetool autorepairstatus inspect and adjust it. Check the documentation for your exact version before relying on any other defaults; the splitter and parallelism settings have their own defaults that are worth reading rather than assuming.

# cassandra.yaml (6.0 line, or 5.0.8+ started with -Dcassandra.autorepair.enable=true)
auto_repair:
  enabled: true                       # default is false
  repair_type_overrides:
    full:
      enabled: true
      min_repair_interval: 5d         # comfortably inside a 10-day gc_grace
    incremental:
      enabled: false                  # turn on only after the migration steps below
-- Opt a table out of incremental auto repair but keep full repair, and raise its priority.
ALTER TABLE orders.line_items
  WITH auto_repair = {'incremental_enabled': 'false', 'full_enabled': 'true', 'priority': '1'};

-- Where the scheduler records what it has done, per node.
SELECT * FROM system_distributed.auto_repair_history LIMIT 20;

Choosing a scheduler

OptionStrengthsWeaknessesChoose it when
Auto Repair (built in)No extra service; topology-aware; history in a system table; per-table opt-outNew; the 5.0.8 switch is one-way; fewer operators have run it at scaleYou are on 6.0 or 5.0.8+ and can validate it in staging first
ReaperMature, UI and API, segment-level retries and pausing, widely runAnother service and backend to operate and upgradeYou already run it, or need its UI and multi-cluster view
cron + nodetoolTransparent, no dependenciesNo coordination, no retries, easy to overlap sessions, silent failuresTiny clusters or short-lived test environments only

The deciding question is not features but ownership: who is paged when repair stops making progress? If the answer is nobody, no scheduler will save you. Pick one scheduler per cluster, never two; overlapping repairs from two tools on the same ranges multiply validation work and can fail each other's sessions.

Full or incremental, per table

Full repair builds Merkle trees over all data in the range every time. It is simple, has no persistent state beyond the SSTables, and its cost scales with data size. Incremental repair only compares data not yet marked repaired, so each run is cheaper, but it splits SSTables into repaired and unrepaired sets, depends on anticompaction, and needs occasional full repair or --validate to catch corruption in the repaired set.

Table shapeRecommendationWhy
Large, append-mostly (events, time series with TWCS)Incremental, plus periodic full or validateMost data never changes; re-validating it every cycle is wasted I/O
Small or medium, heavily updatedFullCheap enough; avoids repaired/unrepaired SSTable split overhead
Delete-heavy with short gc_graceWhichever finishes the cycle fastest, measuredThe deadline dominates every other concern
LWT-heavy tablesEither, but never skip the Paxos step routinelyPaxos state must be repaired too

Whatever you choose, use subrange sessions. A session that covers a large range builds coarse Merkle trees, where one differing partition marks a whole large leaf as different and causes over-streaming; small ranges keep trees fine-grained and failures cheap to retry.

Worked example: does the cycle fit?

Take a 24-node cluster in two datacenters of 12, replication factor 3 in each, about 4 TiB on disk per node, so 48 TiB per datacenter and roughly 16 TiB of distinct data per datacenter. Tables use the default ten-day gc_grace_seconds. Keep three days of margin so one failed cycle can be repeated, leaving a seven-day budget.

The throughput figure must come from your cluster, not a blog: time a few subrange sessions of known size from logs or scheduler history. Suppose they show about 150 GiB per hour per session and the latency SLO tolerates three concurrent sessions cluster-wide. The arithmetic below gives a full cycle of about 36 hours, well inside the 168-hour budget, so full repair every five days is affordable. If the answer were 150 hours, you would reach for incremental repair on the large append-mostly tables, more parallelism if latency allows, or a longer gc_grace_seconds, in that order.

# Does the repair cycle finish inside gc_grace_seconds with room to spare?
GC_GRACE_S = 864_000                 # 10 days, the table default
MARGIN_S   = 3 * 86_400              # keep 3 days for one failed cycle to be re-run

def cycle_ok(total_bytes_per_replica_set, sessions_in_parallel, bytes_per_hour_per_session):
    hours = total_bytes_per_replica_set / (sessions_in_parallel * bytes_per_hour_per_session)
    budget_h = (GC_GRACE_S - MARGIN_S) / 3600
    return hours, budget_h, hours <= budget_h

# Worked example (measured numbers go here, not guesses):
TiB = 1024**4
hours, budget, ok = cycle_ok(
    total_bytes_per_replica_set=48 * TiB / 3,   # 48 TiB on disk per DC, RF 3 -> 16 TiB of distinct ranges
    sessions_in_parallel=3,
    bytes_per_hour_per_session=150 * 1024**3,   # 150 GiB/h, taken from your own session logs
)
print(f"cycle {hours:.0f} h, budget {budget:.0f} h, ok={ok}")   # cycle 36 h, budget 168 h, ok=True

Put the numbers on a dashboard. The single most useful repair metric is coverage age: for each table, the age of the oldest range since its last successful repair. Alert when it passes about 60 percent of gc_grace_seconds and page at 80 percent. That one alert catches a stalled scheduler, a failing table and a cluster that has outgrown its repair budget.

Moving an existing cluster to incremental repair

The Auto Repair documentation carries a warning worth repeating for any tool: enabling incremental repair on a cluster that has never run it can overwhelm nodes with anticompaction, because the first incremental run finds every SSTable unrepaired and splits them. A safe migration:

  1. Complete a full repair cycle of the table so replicas are known to agree.
  2. If your consistency requirements allow it, mark the existing SSTables as repaired with the offline sstablerepairedset tool, node by node with the node stopped, so the first incremental run only sees new data.
  3. Enable incremental repair for one table, watch pending compactions, disk headroom and read latency for a full cycle, then widen.
  4. Keep a periodic full repair or --validate run on a longer interval; incremental repair assumes the repaired set stays correct, and only a full comparison proves it.

Failure modes

  • Stalled scheduler, green dashboards. Repair silently stops (expired credentials, a stuck lock, a node that fails every validation) and nothing alerts until zombie data appears. Coverage age is the defence.
  • Disk exhaustion from streaming. Over-streaming on coarse ranges and anticompaction both need free space. Keep headroom, use small subranges, and watch pending compactions per node during repair windows.
  • Repair overlapping topology changes. Bootstrap, decommission and replacement change range ownership mid-session. Pause repair during topology changes and resume afterwards; streaming during bootstrap already competes for the same bandwidth.
  • Latency spikes from validation. Validation compaction reads every SSTable in the range. Cap concurrency, keep -j low, and schedule the heaviest tables away from peak.
  • Hints masking a lagging node. Hints cover outages shorter than the hint window; hinted handoff is not a substitute for repair and gives no guarantee after the window passes.
  • Rejoining a stale node. A node down longer than gc_grace_seconds rejoining with its old data is the fastest route to resurrected rows. Replace it instead.

Trade-offs in one place

More repair parallelism finishes cycles sooner but costs tail latency. Incremental repair saves I/O but adds state that can itself go wrong. A longer gc_grace_seconds gives repair more slack but keeps tombstones longer, which slows reads over delete-heavy partitions. The built-in scheduler removes a moving part but, on 5.0.8, cannot be turned back off. Each trade-off should be settled with a measurement: cycle time, coverage age, p99 latency during repair and tombstone counts per read. The broader operations model describes how repair fits alongside backups, upgrades and capacity work.

What to do next

  1. List every table with its gc_grace_seconds and last successful repair time; that is your coverage-age baseline.
  2. Time a handful of subrange sessions and compute your full-cycle hours with the arithmetic above.
  3. Add coverage-age alerts at 60 and 80 percent of gc_grace_seconds.
  4. Pick exactly one scheduler; if you are considering Auto Repair on 5.0.8, trial it in staging, remembering the enable flag is one-way.
  5. Classify tables as full or incremental; migrate incremental candidates one at a time, full cycle first.
  6. Run nodetool repair --preview weekly on a sample keyspace and trend the estimated streaming volume.
  7. Write the runbook entry: pause repair for topology changes, and replace rather than rejoin nodes down longer than gc_grace_seconds.
Key takeaway: Repair in 2026 is a deadline problem: every range of every table must be repaired within gc_grace_seconds, with margin. Measure your cycle time, choose full or incremental per table, run one scheduler (the built-in Auto Repair, Reaper or, only for tiny clusters, cron), keep the Paxos step on, and watch coverage age rather than trusting that jobs ran. Move to incremental repair deliberately, after a full cycle, and prove the repaired set periodically with a full or validate run.