Time-series data is the workload Cassandra handles most naturally and still gets wrong most often. Metrics, sensor readings, audit events and price ticks share a shape: rows arrive roughly in time order, they are almost never updated, recent data is read far more than old data, and everything eventually expires. Cassandra's write path, its sorted clustering columns and its time-window compaction suit that shape well. Its one-partition-per-query model punishes a careless design, though, and the punishment usually arrives months after launch, when partitions have grown.

This article builds the model from first principles: bucket arithmetic, reads that span buckets, compaction aligned with expiry, and rollups for dashboards, then failure modes and a checklist. Primary keys in general are covered in Cassandra data modeling basics and the TTL purge rules in Cassandra TTL, in depth; this page applies both to time-ordered data.

Advertisement

Why Cassandra fits time-ordered data

Cassandra is a log-structured merge store: a write goes to the commit log and a memtable, which is later flushed as an immutable SSTable. Nothing is read or rewritten on the write path, so ingest cost barely depends on how much data is stored.

Inside a partition, rows are sorted by clustering columns. With the timestamp as clustering column, a partition is a time-sorted log of one source and a time range is a contiguous slice; with descending order, the latest N points sit at its head.

Expiry fits too: with a table TTL and Time-Window Compaction Strategy (TWCS), a file holding one day's readings eventually contains only expired cells and can be deleted whole. The rest of the design keeps these properties intact.

The reference schema

The partition key is the source identifier plus a time bucket, the clustering key is the timestamp plus a tie-breaker, and the table carries its retention and compaction settings:

CREATE TABLE metrics.readings (
    sensor_id  text,
    day        date,          -- the bucket: part of the partition key
    ts         timestamp,
    seq        int,           -- tie-breaker for readings in the same millisecond
    value      double,
    unit       text,
    PRIMARY KEY ((sensor_id, day), ts, seq)
) WITH CLUSTERING ORDER BY (ts DESC, seq ASC)
  AND default_time_to_live = 2592000          -- 30 days
  AND compaction = {
        'class': 'TimeWindowCompactionStrategy',
        'compaction_window_unit': 'DAYS',
        'compaction_window_size': 1 };

-- the only query shape this table serves:
SELECT ts, value FROM metrics.readings
 WHERE sensor_id = ? AND day = ? AND ts >= ? AND ts < ?;

Each choice has a reason. The day column in the partition key caps how large any partition can grow; without it one sensor's partition would grow forever. ts DESC puts the newest reading first, so a LIMIT 100 for the latest points is cheap. seq exists because two readings can share a millisecond, and writes to the same primary key are upserts, so the second would silently replace the first. The table TTL and a one-day TWCS window match the retention and the bucket.

The table serves one query shape: one sensor, one bucket, a time slice. Fleet-wide questions need their own tables or another system.

Write path and read fan-out for a bucketed time-series tableSensors / agentsappend readingsWriter servicebucket = f(sensor, ts)ts, valuePartition (s-17, 2026-10-01)rows clustered by ts DESCINSERT ... USING TTLPartition (s-17, 2026-09-30)yesterday's bucket, coldPartition (s-17, 2026-09-29)older, expiringDashboard querylast 36 hoursQuery plannerrange to bucket listfrom, toasync 1async 2async 3k-way mergeordered, paged resultTWCS on diskwindow 2026-10-01: SSTables still compactingwindow 2026-09-30: one SSTablewindow 2026-09-29: one SSTablewindow past TTL + gc_grace: whole file dropped(only if no newer data was mixed in)
Writers compute the bucket from the timestamp and append. Readers turn a time range into a list of buckets, query them in parallel and merge. TWCS keeps each day in its own SSTable so expired days can be dropped whole.
Advertisement

Choosing a bucket with arithmetic

The bucket is the most important number in the design: too large and partitions are slow to read, compact and repair; too small and queries fan out. Compute it from the per-source write rate, the row size and the typical query range.

Worked example. Budget about 50 bytes per row on disk before compression for this schema, then measure the real figure with nodetool tablestats:

Write rate per sensorRows per dayRaw bytes per dayBucket that keeps partitions modest
1 reading per minute1,440~70 KBmonth (~2 MB) or even year
1 reading per second86,400~4.3 MBday
10 per second864,000~43 MBday, or 6 hours if queries are short
100 per second8,640,000~430 MBhour (~18 MB)

A widely used rule of thumb keeps partitions in the tens of megabytes and well clear of 100 MB; it is guidance, not a hard limit. The second check is query width: a six-hour dashboard over daily buckets touches one or two partitions, which is ideal, while a week-long chart over hourly buckets fans out across 168. When the two checks disagree, add a rollup table for the wide view rather than compromising the bucket.

Buckets must be computable from the query alone, such as a date, an hour or a YYYY-MM string. A bucket that depends on how many rows already exist forces a lookup before every write.

Reading across buckets

Because the bucket is in the partition key, a range query that crosses a bucket boundary is several queries. Cassandra will not do this for you: an IN on the bucket column is legal, but it makes one coordinator do all the work and return everything at once. Issue the per-bucket queries from the client, in parallel, with token-aware routing so each goes straight to a replica, and merge the ordered streams:

from datetime import datetime, timedelta, timezone
import heapq
from cassandra.cluster import Cluster

session = Cluster(["10.0.0.11", "10.0.0.12"]).connect("metrics")
stmt = session.prepare(
    "SELECT ts, seq, value FROM readings "
    "WHERE sensor_id = ? AND day = ? AND ts >= ? AND ts < ?")
stmt.fetch_size = 2000                       # page size per round trip

def buckets(start, end):
    d = start.date()
    while d <= (end - timedelta(microseconds=1)).date():
        yield d
        d += timedelta(days=1)

def read_range(sensor, start, end):
    futures = [session.execute_async(stmt, (sensor, day, start, end))
               for day in buckets(start, end)]    # one query per bucket, in parallel
    streams = [iter(f.result()) for f in futures] # result() pages lazily as you iterate
    # each bucket is already ordered ts DESC; merge them newest-first
    return heapq.merge(*streams, key=lambda r: (r.ts, -r.seq), reverse=True)

now = datetime.now(timezone.utc)
for row in read_range("s-17", now - timedelta(hours=36), now):
    print(row.ts, row.value)

The prepared statement lets the driver route by token, fetch_size pages each stream so a week of data never sits in memory, and the heap merge costs O(log k) per row. For the latest N points, query the newest bucket first and go further back only if it returned too few rows. Cap the range the API accepts and send wider requests to rollup tables.

Aligning compaction with expiry

TWCS puts SSTables into windows by the maximum write timestamp they contain and compacts only within a window. While a window is current, its SSTables are compacted together using size-tiered rules; once the window closes, its files are compacted down, typically to one, and left alone. When every cell in a file has expired and passed gc_grace_seconds, and no overlapping SSTable could hold data the file's tombstones shadow, the whole file is dropped. Those purge conditions are spelled out in the TTL article; the modeling consequences are below.

  • Size the window against retention. Aim for a few dozen windows across the retention period. Thirty days of retention with one-day windows gives about 30 SSTables per table per node after compaction; a one-hour window would give 720 and slow reads, while a 30-day window would hold the data on disk up to a month longer than needed.
  • Use one TTL per table. If some rows live for 7 days and others for a year, a window cannot be dropped until its longest-lived cell expires. Put different retention classes in different tables.
  • Do not update or delete old rows. A late write or an explicit delete lands in the current window but refers to old data, so the old window and the new file now overlap and neither can be dropped cleanly. If backfill is unavoidable, write it through a separate path and expect that disk usage will drop later than planned.
  • Watch repair and read repair. Streaming during repair can bring old cells into new SSTables with old timestamps, mixing windows. Many teams lower or disable read repair chance on TWCS tables (on versions where it is a table option) and schedule repair carefully for the same reason.

Latest values and rollups

Two read patterns dominate time-series front ends: the current value of each source, and charts over long ranges. Neither should be served from the raw table.

-- latest value: one row per sensor, overwritten on every reading (an upsert, no read-before-write)
CREATE TABLE metrics.latest (
    sensor_id text PRIMARY KEY,
    ts        timestamp,
    value     double
);

-- 1-minute rollups written by a stream job, kept far longer than raw data
CREATE TABLE metrics.rollup_1m (
    sensor_id text,
    month     text,             -- '2026-10': ~44,640 rows per sensor per month
    minute    timestamp,
    n         int,
    sum       double,
    min       double,
    max       double,
    PRIMARY KEY ((sensor_id, month), minute)
) WITH CLUSTERING ORDER BY (minute DESC)
  AND default_time_to_live = 31536000        -- 365 days
  AND compaction = {
        'class': 'TimeWindowCompactionStrategy',
        'compaction_window_unit': 'DAYS',
        'compaction_window_size': 14 };

The latest table is a plain upsert where the highest write timestamp wins, so a delayed reading can overwrite a newer one. Write it with USING TIMESTAMP set from the reading's own time in microseconds so the newest reading wins regardless of arrival order.

Rollups are computed outside Cassandra, typically by a stream processor that aggregates one minute at a time and writes count, sum, min and max. Store sum and count rather than an average, so minutes can be combined into hours exactly. Counters are a poor fit: increments are not idempotent, so a retried write double-counts.

Each resolution gets its own table, bucket and TTL, for example raw data for days, minutes for months and hours for years. The read API picks the coarsest table that still gives enough points for the chart width.

Failure modes

SymptomCauseFix
Partitions grow without limit; reads and compaction slow down over monthsNo bucket in the partition keyAdd a time bucket; migrate by writing new data to a new table
One node hot while others idlePartition key is only the bucket (all sensors write to 'today')Put the source id in the partition key; add a shard suffix for single very hot sources
Readings silently missingTwo readings in the same millisecond share a primary key and overwriteAdd a sequence or use a timeuuid clustering column
Disk does not shrink after TTLMixed TTLs, deletes, late writes or repair mixed windowsOne TTL per table, no deletes, check overlapping SSTables with sstablemetadata
Tombstone warnings on readsRange deletes or explicit nulls written as cellsNever write nulls; let TTL expire data instead of deleting; see the tombstones article
Coordinator timeouts on wide chartsLarge IN queries or unpaged reads across many bucketsParallel per-bucket queries with paging; rollups for long ranges

Binding null for unset columns writes a tombstone on every insert; leave the parameter unset instead. ALLOW FILTERING on a time-series table usually means a missing table. The mechanics are in Cassandra tombstones.

Operating a time-series cluster

A few measurements tell you whether the model is holding up. nodetool tablehistograms metrics.readings shows the partition size distribution, which should match your bucket arithmetic. nodetool tablestats gives SSTable counts and tombstones scanned per read, nodetool compactionstats shows whether compaction keeps up, and sstablemetadata reports each file's timestamp range, which finds the SSTable blocking a window drop.

Consistency is a business decision. Raw metrics are often written at LOCAL_ONE because a lost reading is tolerable and latency is not; billing events and audit logs need LOCAL_QUORUM on both writes and reads. Whatever you choose, keep the coordinator token-aware and keep batches single-partition: an unlogged batch of readings for one sensor and one bucket is efficient, while a batch across many partitions only moves the fan-out work onto one coordinator.

Capacity is rows per second times bytes per row times retention times replication factor, divided by the compression ratio, plus compaction headroom. How partitions map to replicas is covered in Cassandra partitioning.

Trade-offs against a dedicated time-series database

Cassandra gives you linear write scaling, multi-datacenter replication, tunable consistency and predictable per-partition reads. It does not give you server-side downsampling, gap filling or aggregation across partitions. For ad hoc analytics over many series, a purpose-built time-series store or a columnar warehouse is the better tool; for high-volume ingest with known read paths, such as per-device history or latest state, Cassandra with this model is hard to beat. The compaction strategies overview compares the alternatives to TWCS if your data does not expire uniformly.

What to do next

  1. Write down every query the product needs, with its time range and acceptable latency, before creating any table.
  2. Measure or estimate the per-source write rate and row size, then compute rows and bytes per bucket for at least three candidate buckets.
  3. Pick the bucket that keeps partitions in the tens of megabytes and keeps the typical query to one or two partitions; add rollup tables for longer ranges.
  4. Add a tie-breaker clustering column so same-millisecond readings cannot overwrite each other.
  5. Set default_time_to_live and TWCS with a window giving a few dozen windows over retention; keep one TTL per table.
  6. Implement the reader as parallel per-bucket prepared queries with paging and a k-way merge, and cap the range the API accepts.
  7. Stop writing nulls, forbid deletes and updates on the raw table, and route backfills through a separate path.
  8. After a week of production traffic, check tablehistograms and tablestats against your arithmetic, then confirm a day after retention ends that expired SSTables are actually being dropped.
Key takeaway: A Cassandra time-series model is a partition key of source plus a computable time bucket, a descending timestamp clustering key with a tie-breaker, one table TTL and TWCS windows sized against retention. Choose the bucket from write-rate arithmetic and query width, read across buckets with parallel paged queries and a merge, serve long ranges from rollup tables, and keep deletes, nulls and late writes away from the raw table so expired days can be dropped as whole files.