Time-to-live is the one deletion mechanism in Cassandra that application code can use without writing a delete. You attach a lifetime to a write, and after it passes the value stops being returned. That makes TTL the natural tool for sessions, caches, leases, idempotency windows and retention policies, and it is also how teams accidentally build tables that never expire, rows that vanish early, or retention changes that cannot be applied to data already on disk.

This article treats TTL as an application design tool. It covers setting TTLs safely from driver code, choosing between TTL, delete and dropping whole tables, five patterns with working CQL and Python, how to change retention for data that is already written, how to audit what is really expiring, and what TTL does not promise. Storage internals, such as liveness markers, the purge rule and time-window compaction, are covered in Cassandra TTL internals; they appear here only where they change an application decision.

Patterns in one picture

TTL patterns and the table design each one needsLogin servicesessionsSchedulersleasesIngest APIidempotency keysTelemetryper-tier retentionINSERT ... USING TTL ?full row, refreshedIF NOT EXISTS + TTLrenew with IF ownerINSERT ... TTL 24hone row per keyone table per tieruniform default TTLLCS or STCSpoint reads, overwritesPaxos per partitionserial consistencySTCS, short gracewrite onceTWCS per tablewhole-file dropsMixing lifetimes in one table is the root of most TTL operational pain.
Each common TTL use maps to a write shape and a table design: refreshed full rows for sessions, conditional writes for leases, write-once keys for idempotency, and a table per retention tier so time-window compaction can drop whole files.

Setting TTLs from code without disabling them

A TTL comes from one of two places: the table's default_time_to_live, applied to any write that does not set its own, or a per-statement USING TTL. With prepared statements the TTL can be a bind marker, so one statement serves many lifetimes. The trap is that a TTL of 0, and a null bound to the marker, both mean no expiry, overriding the table default. Validate in code, at one choke point:

from cassandra.cluster import Cluster

session = Cluster(["10.0.0.1"]).connect("auth")
PUT = session.prepare(
    "INSERT INTO sessions (sid, user_id, device, created_at) "
    "VALUES (?, ?, ?, ?) USING TTL ?")

MAX_TTL = 630_720_000                      # 20 years, Cassandra's upper bound

def put_session(sid, user_id, device, created_at, ttl):
    if not isinstance(ttl, int) or not 0 < ttl <= MAX_TTL:
        raise ValueError(f"refusing ttl={ttl!r}: 0 or None would disable expiry")
    session.execute(PUT, (sid, user_id, device, created_at, ttl))

Use INSERT with every column for anything that must expire as a unit. An UPDATE writes only the columns it names and no row marker, so rows assembled from updates expire column by column; the internals article explains why. Counter tables cannot carry TTLs at all, so anything counted and expired, such as rate limits, needs time-bucketed rows or a different store.

TTL, delete or drop the table?

MechanismBest forCost
TTL on each writeData whose lifetime is known when writtenExpired cells linger until compaction; lifetime fixed per write
Explicit DELETEData whose end is decided later, by an eventA tombstone per delete, read-path cost, repair discipline
Time-bucketed tables, dropped or truncatedLarge uniform datasets with a clear cut-offApplication routes writes and reads by bucket; schema churn
TTL plus TWCSTime series with one uniform TTLBreaks if old timestamps or mixed TTLs enter a window

The question to ask is when the lifetime becomes known. If it is known at write time, use TTL. If it depends on a later event, such as an account closure, use a delete, or design so the delete removes a whole partition. If the dataset is large and uniform, dropping a bucketed table is the cheapest deletion Cassandra offers.

Pattern: sliding-expiration sessions

A web session should expire after thirty minutes of inactivity, not thirty minutes after login. A TTL cannot be extended in place; you extend it by writing the row again. Rewriting on every request multiplies write volume, so refresh only when the remaining TTL has dropped below half the window, which TTL() tells you on the read you already make:

GET = session.prepare(
    "SELECT user_id, device, created_at, TTL(device) AS left FROM sessions WHERE sid = ?")
IDLE = 1800                                 # 30-minute idle timeout

def touch(sid):
    row = session.execute(GET, (sid,)).one()
    if row is None:
        return None                         # expired, or never existed
    if row.left is None or row.left < IDLE // 2:
        put_session(sid, row.user_id, row.device, row.created_at, IDLE)
    return row

TTL() returns null for a column with no TTL or a null value, so pick a column that is always set. The refresh rule bounds the idle timeout to between fifteen and thirty minutes of real inactivity, and cuts writes several-fold for chatty clients. Because each refresh overwrites the same partition, choose leveled compaction or the default size-tiered strategy, not time-window compaction, which assumes data is written once.

Pattern: leases with lightweight transactions

A lease lets one worker own a job until it stops renewing. Lightweight transactions give the compare-and-set and TTL gives the automatic release:

CREATE TABLE ops.leases (name text PRIMARY KEY, owner text, epoch bigint);

-- acquire: succeeds only if no live row exists ([applied] = true)
INSERT INTO ops.leases (name, owner, epoch) VALUES ('compactor', 'host-17', 42)
IF NOT EXISTS USING TTL 30;

-- renew every ~10 s: rewrite both columns with a fresh TTL, only while still owner
UPDATE ops.leases USING TTL 30 SET owner = 'host-17', epoch = 42
WHERE name = 'compactor' IF owner = 'host-17';

-- release early on clean shutdown
DELETE FROM ops.leases WHERE name = 'compactor' IF owner = 'host-17';

Once the owner stops renewing, both cells expire and the next IF NOT EXISTS succeeds. Renew at a third of the TTL so one slow renewal does not lose the lease. A lease is not a lock that protects data by itself: a worker paused by garbage collection can wake after its lease expired and still act. Give each acquisition a strictly increasing epoch, taken from a separate non-expiring row incremented with a conditional update, and have the protected resource reject writes carrying an older epoch. Lightweight transactions cost several round trips each; see lightweight transactions before using them on a hot path.

Pattern: idempotency windows and caches

An ingest API that accepts retries needs to remember idempotency keys for as long as clients may retry, and no longer. One row per key with a fixed TTL does it: INSERT INTO ingest.seen (key, result) VALUES (?, ?) USING TTL 86400. If two concurrent retries must not both run, the insert needs IF NOT EXISTS, and the applied flag decides which proceeds. If an occasional duplicate is harmless, skip the transaction and read before write. A read-through cache is the same shape: write the computed value with a TTL equal to how stale it may be, and treat a miss as a cue to recompute. Both are write-once tables, so expired data is purgeable at its first compaction once the write is older than gc_grace_seconds; for short TTLs on tables that never see explicit deletes, a shorter grace period reclaims space sooner, provided hints and repair complete within it.

Pattern: retention tiers

Multi-tenant products often sell retention: seven days on one plan, ninety on another. Passing a per-tenant TTL into one shared table works for correctness and fails for operations, because a time-window table full of mixed lifetimes can never drop a window whole while any long-lived row remains in it, and size-tiered tables mix ages in every file. Put each retention tier in its own table with its own default_time_to_live and compaction window, and route by tier in the data-access layer. A tenant who upgrades gets new writes in the longer table, and reads query both until the old data ages out.

Changing the retention of data already written

Altering default_time_to_live affects only future writes; every existing cell keeps the expiry it was written with. To shorten or lengthen the life of existing data you rewrite it. The trick is to rewrite each row with a write timestamp one microsecond above its original, so the rewrite supersedes the old version but any application write made since the scan, carrying a current timestamp, still wins. Use INSERT, not UPDATE, so the row marker is rewritten with the new TTL too:

# retention_rewrite.py - make rows expire NEW_TTL after their ORIGINAL write.
import time
from cassandra.query import SimpleStatement

NEW_TTL = 7 * 86400
scan = SimpleStatement(
    "SELECT sensor_id, day, ts, value, WRITETIME(value) AS wt FROM metrics.readings",
    fetch_size=1000)                         # better: split by token range, run in parallel
rewrite = session.prepare(
    "INSERT INTO metrics.readings (sensor_id, day, ts, value) VALUES (?, ?, ?, ?) "
    "USING TTL ? AND TIMESTAMP ?")
drop = session.prepare(
    "DELETE FROM metrics.readings USING TIMESTAMP ? WHERE sensor_id = ? AND day = ? AND ts = ?")

for r in session.execute(scan):
    left = int(NEW_TTL - (time.time() - r.wt / 1_000_000))
    if left > 0:
        session.execute(rewrite, (r.sensor_id, r.day, r.ts, r.value, left, r.wt + 1))
    else:
        session.execute(drop, (r.wt + 1, r.sensor_id, r.day, r.ts))

Rows with several regular columns may carry different write times per column; rewrite from the newest, or rewrite each column separately. Shortening retention this way creates a burst of tombstones and rewritten cells, so throttle it and watch compaction. For very large tables, writing the surviving window into a new table and switching reads over is often cheaper than rewriting in place.

Auditing what really expires

Do not assume a table expires because its schema has a default. Sample it:

rows = session.execute(SimpleStatement(
    "SELECT TTL(device) AS t FROM auth.sessions LIMIT 20000", fetch_size=1000))
ttls = [r.t for r in rows]
missing = sum(t is None for t in ttls)
print(f"{missing}/{len(ttls)} sampled rows will never expire")
print("longest remaining:", max((t for t in ttls if t is not None), default=None))

An unbounded LIMIT scan walks token order, so it is a biased but cheap sample; repeat it with token-range predicates for coverage. Any non-zero missing count means some writer is sending 0 or null, writing without a TTL on a table without a default, or using UPDATE paths. Track disk usage on TTL tables over weeks: at steady state it should be flat.

What TTL does not guarantee

  • It is not erasure on a deadline. Expired data stays in SSTables until compaction purges it, and it survives in snapshots, incremental backups and any copies made before expiry. If a regulation requires deletion within a fixed time, design for it explicitly: shorter grace periods, scheduled compaction, backup retention, or per-user encryption keys you can destroy.
  • It is judged by server clocks. Expiry is evaluated against the nodes' clocks, so skew between nodes shifts the instant a value disappears by that skew. Keep NTP healthy and never use TTL boundaries for sub-second logic.
  • It is not free to read past. Expired cells still count against tombstone thresholds when a query scans over them, as in queue-like tables read from the oldest end.
  • It caps at twenty years, and long TTLs interact with the 2038 limit on clusters still in Cassandra 4 storage compatibility mode; the internals article has the details.

Worked example: sizing a session table

A login service has 3 million sessions created a day, a 30-minute idle timeout and about 20 authenticated requests per session over a typical 40-minute visit. Rewriting on every request would mean about 60 million writes a day. With the half-window refresh rule a session is rewritten at most every 15 minutes, about three times per visit, so writes fall to roughly 12 million a day including the initial insert.

Live data is tiny: even 500,000 concurrent sessions at 300 bytes are 150 MB. Dead data comes in two kinds. Superseded versions, about 9 million refreshes or 2.7 GB a day before replication, are discarded as soon as compaction merges them with the newer version of the same row; how long that takes depends on the compaction strategy, which is the real reason to prefer leveled compaction here. Each session's final version is different: once expired it waits until its write time plus gc_grace_seconds before it can be purged. That is 3 million rows, about 0.9 GB a day, so roughly 9 GB per replica set sits on disk under the default ten-day grace. The table never receives explicit deletes and repair runs daily, so lowering grace to one day cuts that part to about 0.9 GB. Every one of those numbers came from a design decision, not from Cassandra.

Failure modes

  • Retention silently off. An ORM binds null to USING TTL ?; rows live forever. Validate TTLs at one choke point and audit by sampling.
  • Sessions that never time out. A profile update uses UPDATE with no TTL on one column; the row survives its expiry. Rewrite whole rows.
  • Leases held by the dead. TTL longer than the failure-detection budget delays failover; TTL too short loses leases to pauses. Renew at a third of TTL and fence with epochs.
  • Retention changes that do nothing. Altering the default leaves existing data untouched. Rewrite, or migrate to a new table.
  • Disk that keeps growing. Overwritten versions waiting on a compaction strategy that rarely merges them, expired rows waiting out a long grace period, or mixed lifetimes in time-window tables. Pick compaction for the write pattern, shorten grace where safe and separate tiers.
  • Tombstone warnings on reads. Scans over expired regions of wide partitions. Bucket partitions by time and read only live buckets.

What to do next

  1. Route every TTL'd write through one function that rejects 0, null and values above 630,720,000.
  2. Sample TTL() on every table that is supposed to expire and fix every writer that produces rows without one.
  3. Replace column UPDATE paths on expiring rows with full-row INSERT statements.
  4. Move per-tenant or per-plan retention into one table per tier, each with its own default TTL and compaction settings.
  5. For sliding sessions, add the half-window refresh rule and measure the write reduction.
  6. For each TTL-only table, compare TTL, grace period and repair cadence, and shorten grace deliberately where it is safe; read Cassandra tombstones first.
  7. Write down where TTL is not enough, such as legal erasure or backups, and design those paths explicitly.
Key takeaway: TTL is the right Cassandra deletion tool when a value's lifetime is known at write time. Set it through one validated code path so 0 or null never disables expiry, write expiring rows as full INSERTs, refresh sliding sessions only past half their window, pair leases with conditional writes and fencing epochs, keep each retention tier in its own table, and rewrite existing data with its original write time plus one microsecond when retention changes. Audit expiry by sampling TTL(), and remember TTL is neither prompt erasure nor free to read past.