A time-to-live (TTL) tells Cassandra to treat a value as deleted a fixed number of seconds after it was written. It is the standard way to keep sessions, caches, time series and audit trails from growing forever, and it looks like the cheapest kind of deletion there is: no delete statement, no batch job, no application logic. Underneath, an expired value goes through the same machinery as an explicit delete, and whether its disk space ever comes back depends on how you modelled the table and which compaction strategy you chose.

This article explains TTL from the storage engine up: what is stored with each cell, how rows stay alive or disappear, what the table default really does, how an expired cell becomes a tombstone and when it can be purged, and why time-window compaction makes TTL cheap only when the data cooperates. General tombstone behaviour, zombie data and query thresholds are covered in Cassandra tombstones; this page stays on TTL.

Advertisement

TTL lives on the cell, not the row

Cassandra stores data as cells: one value for one column of one row, carrying a write timestamp in microseconds. A cell written with a TTL also carries the TTL in seconds and a local expiration time, the moment it stops being live, computed as the time of the write plus the TTL. Every replica receives the same TTL and expiration with the mutation, which is why expiry needs no coordination: each replica independently stops returning the cell when its clock passes the expiration time.

Because the TTL belongs to the cell, two columns of one row can expire at different times, and so can two elements of the same collection. The CQL documentation states it plainly: the TTL concerns the inserted values, not the columns, and any later update of a column resets its TTL to whatever that update specifies. You can see both properties with the TTL() and WRITETIME() functions, which cannot be applied to primary-key columns.

Row liveness: the INSERT versus UPDATE trap

A row exists in a query result if it has at least one live cell or a live primary-key liveness marker. An INSERT writes that marker, and the marker gets the statement's TTL. An UPDATE does not write the marker; it writes only the columns it sets. The difference changes what disappears when.

CREATE TABLE app.sessions (
  user_id  uuid,
  sid      text,
  device   text,
  token    text,
  PRIMARY KEY ((user_id), sid)
) WITH default_time_to_live = 86400;            -- 1 day for writes that do not say otherwise

-- Case 1: INSERT with the default TTL, then UPDATE one column with no TTL.
INSERT INTO app.sessions (user_id, sid, device, token) VALUES (?, 's1', 'ios', 't0');
UPDATE app.sessions USING TTL 0 SET device = 'ios-17' WHERE user_id = ? AND sid = 's1';
-- After a day: token and the liveness marker expire, device does not.
-- The row still appears, with token = null. Sessions "never expire".

-- Case 2: UPDATE-only rows.
UPDATE app.sessions USING TTL 600 SET token = 't1' WHERE user_id = ? AND sid = 's2';
-- No liveness marker was written. When token expires, the row vanishes entirely,
-- even though no one deleted it.

SELECT sid, device, TTL(device), token, TTL(token), WRITETIME(token)
FROM app.sessions WHERE user_id = ?;

The fix is a modelling rule: write rows that must expire as a unit with a single INSERT carrying all columns and one TTL, and if a column later changes, rewrite the whole row with a fresh TTL rather than updating one column. Treat UPDATE ... USING TTL as a per-column tool, not a row tool.

Advertisement

default_time_to_live, TTL 0 and changing TTLs

The table option default_time_to_live is applied, in seconds, to any write that does not specify its own TTL. It is evaluated at write time and stored in the cells, so changing it with ALTER TABLE affects only future writes; existing data keeps the TTL it was written with. To shorten the lifetime of data already on disk you must rewrite it or delete it.

A per-statement USING TTL overrides the default. According to the CQL documentation, a TTL of 0 is equivalent to no TTL, and on a table with a default it removes the TTL for the values written; a TTL of null behaves like 0. That is useful for pinning an individual row, and dangerous when a driver or ORM sends 0 or null for "unspecified", silently turning the table's retention policy off for those writes.

Two limits apply. The maximum TTL is 20 years (630,720,000 seconds). Counter tables do not support TTL at all, so expiring counters means time-bucketed tables you drop or truncate.

Reading expired data

Expiry is lazy. When a cell's expiration time passes, nothing happens on disk. The cell stays in its SSTable until a compaction rewrites that file. Reads see it, compare its expiration with the current time, and drop it from the result, as part of the normal merge across memtable and SSTables described in the write and read path.

That filtering is not free. Expired cells are counted with tombstones for the purposes of tombstone_warn_threshold and tombstone_failure_threshold, whose defaults are 1,000 and 100,000 per query. A partition that accumulates many TTL'd rows and is read with a range scan over the expired part will log warnings and eventually fail queries, even though the application never issued a delete. The classic victim is a queue-like table where consumers always read from the oldest end.

From expired cell to purged: the compaction rule

When compaction reads an expired cell, it decides between dropping it and keeping it as a tombstone. The logic, in the cell's purge method in the Cassandra source, works like this: an expired cell whose expiration time is already older than gc_grace_seconds ago may be dropped. Otherwise it is rewritten as a tombstone whose local deletion time is the expiration time minus the TTL, which is the original write time, and that tombstone is checked again. So the effective rule is that an expired cell can be purged once its write time plus gc_grace_seconds has passed.

The life of a TTL'd cell: live, expired-but-present, tombstone, purgedLive celltimestamp, ttl, expiryExpired cellfiltered at read timeTombstoneafter compaction rewrites itPurgedgone from disknow >= expirycompaction, not yet purgeablecell: write time + gc_grace passedno overlapping older dataTWCS: one SSTable per time windowday 1all expiredday 6mixedday 7liveDropped whole once newest expiry + gc_gracehas passed, without reading it.One old row in a new window blocks the drop.Read pathMerge cells across memtable + SSTablesDrop expired cells; count them as tombstones
A TTL'd cell is filtered at read time once expired, rewritten as a tombstone by compaction, and purged when the grace period measured from its write time has passed and no overlapping SSTable holds older data for the same key. TWCS drops a whole SSTable once its newest expiry plus the grace period has passed.

Two consequences follow. With a TTL at least as long as gc_grace_seconds (default 864,000 seconds, ten days), expired cells are usually purgeable at the first compaction after expiry, because every replica received the TTL with the original write and there is nothing to propagate. With a short TTL and the default grace period, expired data lingers as tombstones for up to ten days after the write. For TTL-only tables with no explicit deletes, some teams lower gc_grace_seconds to shorten that window, but it also bounds how long hints and repair can lag, so do it only when repair cadence allows; see Cassandra repair.

The second gate is overlap. Compaction may only purge a tombstone if no SSTable outside the compaction holds older data for the same partition, otherwise that older data would reappear. Expired data sharing partitions with data spread across many SSTables can therefore sit on disk long after both conditions on time are met.

TWCS: dropping whole files

Time-window compaction strategy groups SSTables by the time window of their writes, compacts within each window, and stops touching a window once it is closed. Whole-file expiry uses its own rule: an SSTable whose newest expiration time is older than gc_grace_seconds ago, and whose data is all older than any overlapping live data, is dropped without being read. Note that this waits for expiry plus grace, not write time plus grace, which is why TWCS tables often run a shorter grace period. For time series with a uniform TTL, that turns retention into deleting a few files per window, which is as cheap as deletion gets. The trade-offs among strategies are compared in compaction strategies.

CREATE TABLE metrics.readings (
  sensor_id  text,
  day        date,
  ts         timestamp,
  value      double,
  PRIMARY KEY ((sensor_id, day), ts)
) WITH default_time_to_live = 2592000           -- 30 days
  AND gc_grace_seconds = 10800                  -- only if TTL-only and repair keeps up
  AND compaction = {'class': 'TimeWindowCompactionStrategy',
                    'compaction_window_unit': 'DAYS',
                    'compaction_window_size': 1};

# Which SSTables are blocking a fully expired one from being dropped?
sstableexpiredblockers metrics readings

Aim for roughly 20 to 30 windows over the TTL, so a 30-day TTL with one-day windows is a reasonable fit. The drop fails, and the files pile up, whenever one SSTable contains data that is not expired: an old row backfilled with an old timestamp into a new window, a row written without the table's TTL, a repair or read repair that streams older data into a recent file, or explicit deletes mixed in. The sstableexpiredblockers tool names the SSTables preventing a drop. On Cassandra 4.0 and later, setting the table's read_repair option to 'NONE' stops read repair from mixing old data into new windows, at the cost of the monotonic-read guarantee that blocking read repair provides.

The 2038 and 2106 limits

The local expiration time was historically stored as a 32-bit count of seconds, so no cell could expire after 2038-01-19T03:14:06 UTC. As that date approaches, long TTLs collide with it: a 20-year TTL written today already exceeds it. Cassandra's NEWS file documents that by default such writes are rejected, and that other overflow policies can be chosen; check the policies available for your exact version before changing it, because capping a TTL silently shortens retention.

Cassandra 5.0 raised the representable maximum to 2106-02-07T06:28:13 UTC, but only after the cluster leaves compatibility mode. The storage_compatibility_mode setting defaults to CASSANDRA_4, which keeps the 2038 limit and the ability to roll back. The documented path is a rolling upgrade to 5.0, a rolling restart with UPGRADING, and a final rolling restart with NONE. If you store multi-year TTLs, plan that sequence early.

Worked example: a 30-day sensor table

A fleet of 50,000 sensors writes one reading per minute with a 30-day TTL into the table above. At steady state each daily partition has 1,440 rows, the cluster holds 30 windows, and every day one window's SSTables become fully expired and are dropped. Disk usage is flat.

Then an operator backfills a week of readings recovered from a gateway, with original timestamps, so they land in today's window. Today's SSTables now contain data that will expire in about 23 days while the rest of the file expires in 30, and worse, the backfilled rows overlap partitions in older windows. Fully expired files start reporting blockers, and disk grows by a day's worth each day. The repair is to stop writing old timestamps into current windows: backfill into a separate table, or accept that those rows are rewritten by a major compaction of the affected windows. Measure with nodetool tablestats and sstableexpiredblockers before and after.

Failure modes

  • Rows that never expire. A later UPDATE with TTL 0 or null, or without a TTL on a table with no default, keeps the row alive after the original columns expire.
  • Rows that vanish. Update-only rows have no liveness marker and disappear when their last TTL'd column expires.
  • Retention silently off. Drivers sending TTL 0 or null override default_time_to_live.
  • Tombstone-threshold failures. Range scans over expired regions of a partition hit the same limits as explicit deletes.
  • Files that never drop. Backfills, mixed TTLs, read repair and deletes block TWCS whole-SSTable expiry.
  • Year-2038 rejections. Long TTLs start failing writes on clusters still in 4.x compatibility mode.

Operations and trade-offs

Watch tombstone-scanned histograms per table, the estimated droppable tombstones that sstablemetadata reports, SSTable count per table and disk growth on TTL'd tables, which should be flat at steady state. TTL is ideal for uniformly ageing data written once; it is a poor fit for data whose lifetime changes after writing, because changing a TTL means rewriting every cell. For whole partitions with a known end, time-bucketed tables that you truncate or drop can be cheaper still.

What to do next

  1. List every TTL'd table and check that rows are written by one INSERT with all columns, not by column updates.
  2. Grep your data access layer for TTL 0 or null being sent by default.
  3. Compare each table's TTL with its gc_grace_seconds and repair cadence, and decide deliberately whether to lower grace.
  4. Move uniformly expiring time series to TWCS with 20 to 30 windows per TTL, and keep backfills out of current windows.
  5. Run sstableexpiredblockers on TWCS tables whose disk is growing.
  6. If you use TTLs longer than a decade, plan the Cassandra 5.0 storage-compatibility sequence.
Key takeaway: A Cassandra TTL is stored on each cell, so rows expire column by column, and only INSERT gives a row a liveness marker. Expired cells are filtered at read time and count against tombstone thresholds. Compaction can purge them once the write time plus gc_grace_seconds has passed and no overlapping SSTable holds older data. TWCS makes expiry nearly free by dropping whole files, but only if each window contains uniformly expiring data. Model rows as single INSERTs, guard against TTL 0, and keep old data out of new windows.