Apache Cassandra 5.0 is the largest change to the storage engine and query surface in years. It adds a new memtable and a new SSTable format built on tries, a compaction strategy that subsumes the old ones, storage-attached indexes and vector search, dynamic data masking, new CQL functions, a later TTL expiry limit, and a Java baseline of 11 with full Java 17 support. Most of it is opt-in, which is good for upgrades and bad for anyone who assumes that installing 5.0 turns everything on.

This article explains each feature in terms of the write and read path it changes, what it costs, and when to enable it, then walks through upgrading a 4.1 cluster and adopting the new defaults safely. Facts come from the Cassandra 5.0 documentation and NEWS file as read on 2026-10-03; release dates, patch versions and benchmark numbers are deliberately left out.

Where the features act

Client (CQL)masking, new functionswriteMemtableskiplist or trie (CEP-19)flushSSTablebig or bti format (CEP-25)SAI index componentsper SSTable (CEP-7, CEP-30)built at flushcompactionUnified Compaction StrategyLf, Tf, N, sharded (CEP-26)storage_compatibility_modeCASSANDRA_4, UPGRADING, NONEgates new formatsJVMJava 11 or 17SecurityCIDR authorizer, crypto providerread
5.0 features mapped onto the write path, flush, compaction and query; storage_compatibility_mode decides when nodes may write the new formats.

The feature list

The documented 5.0 feature list, grouped by where it acts:

AreaFeatureTrackingOn by default after upgrade?
MemtableTrie memtableCEP-19No: configure it
SSTableTrie-indexed SSTables (BTI)CEP-25No: select the format
CompactionUnified Compaction StrategyCEP-26No: per table or default_compaction
IndexingStorage-attached indexes (SAI)CEP-7Available; you create indexes
QueryVector type and similarity functionsCEP-30Available
SecurityDynamic data maskingCEP-20No: dynamic_data_masking_enabled
SecurityCIDR authorizerCEP-33No
CQLMath functions, collection functions, WRITETIME and TTL on collections and UDTsCASSANDRA-17221, 18060, 8877Available
StorageExpiration limit moved from 2038 to 2106CASSANDRA-14227Only after the compatibility steps
RuntimeJDK 17 supported, Java 8 removedCASSANDRA-16895Yes
Toolssstablepartitions, a system-logs virtual table, Azure snitch, pluggable crypto provider, more guardrailsvariousAvailable

Storage engine: trie memtables and BTI SSTables

Trie memtables. The classic memtable is a concurrent skip list of partitions, with each partition holding its own on-heap structures. Under heavy writes that means many small, long-lived objects, which is exactly what the garbage collector handles worst. The trie memtable stores partitions in an in-memory trie whose indexing structure lives in a buffer that can be off-heap, which the documentation credits with better GC behaviour, memory efficiency and lookup speed. Memtable implementations are selected through named configurations in cassandra.yaml and referenced per table:

# cassandra.yaml (example)
memtable:
  configurations:
    skiplist:
      class_name: SkipListMemtable
    trie:
      class_name: TrieMemtable
    default:
      inherits: skiplist     # change to trie to switch every table without its own setting

-- CQL: opt one table in or out by configuration name
ALTER TABLE shop.orders WITH memtable = 'trie';
ALTER TABLE shop.orders WITH memtable = 'default';

BTI SSTables. The traditional ("big") format finds a partition through a sampled index summary held in memory, then a partition index on disk, often helped by the key cache. The BTI format replaces those with trie-structured partition and row indexes that are more compact and cheaper to search, which is why the cassandra_latest.yaml template ships with the key cache set to 0. The format is chosen with sstable: selected_format: bti. Older SSTables stay readable and are rewritten into the new format by normal compaction or by nodetool upgradesstables. Earlier versions cannot read BTI files, so writing them is a point of no return for rollback.

Unified Compaction Strategy

Before 5.0 you picked between size-tiered (STCS: cheap writes, more SSTables per read), leveled (LCS: few SSTables per read, heavy write amplification) and time-window (TWCS) compaction, and switching meant a large recompaction. The Unified Compaction Strategy expresses all of these as positions on one scale, controlled by scaling_parameters:

  • Tf, for example T4: tiered, favouring writes; merge when about f similar-sized SSTables accumulate.
  • Lf, for example L10: leveled, favouring reads; each level holds about f times the previous one.
  • N: the middle point, fan-out 2, behaving like both.

UCS also shards each level by token range (base_shard_count defaults to 4, and target_sstable_size to 1 GiB), so compactions on different shards run in parallel and, unlike LCS, a compaction in one level does not block others. Its documentation states that parameters can change in flight to move between behaviours without a full recompaction.

-- Write-heavy event table: start tiered
ALTER TABLE shop.events WITH compaction = {
  'class': 'UnifiedCompactionStrategy', 'scaling_parameters': 'T4' };

-- Read-heavy lookup table: leveled behaviour
ALTER TABLE shop.orders WITH compaction = {
  'class': 'UnifiedCompactionStrategy', 'scaling_parameters': 'L10' };

Leave TWCS in place for genuine time-series tables with TTLs until you have measured UCS on a copy; TWCS's whole-window expiry is a different optimisation. DateTieredCompactionStrategy is gone in 5.0: move any such table to TWCS before upgrading, or the node will not handle it.

SAI and vector search

SAI attaches index components to each SSTable and memtable instead of keeping a separate hidden table, so indexes are built at flush, merged by compaction and streamed with the data. It supports equality, range and multi-column queries and is the index type to use in 5.0; the cassandra_latest.yaml template makes sai the default for CREATE INDEX. The internals are covered in storage-attached indexing.

Vectors add a vector<float, n> type, an SAI-backed approximate nearest-neighbour index and similarity functions, so ORDER BY embedding ANN OF [...] LIMIT k works next to ordinary columns. Cost, sizing and failure modes are covered in Cassandra 5 vector search.

Dynamic data masking

Dynamic data masking lets the schema decide what unprivileged readers see, without copying data into a redacted table. It is off by default; set dynamic_data_masking_enabled: true in cassandra.yaml. A mask is a function attached to a non-primary-key column. Readers without the UNMASK permission get the function's output; the stored data is unchanged.

CREATE TABLE clinic.patients (
    id     timeuuid PRIMARY KEY,
    name   text MASKED WITH mask_inner(1, null),   -- keep the first character, pad the rest with *
    phone  text MASKED WITH mask_outer(0, 4),      -- pad the last four characters
    birth  date MASKED WITH mask_default()         -- type default: '****', 0, false, ...
);

ALTER TABLE clinic.patients ALTER name MASKED WITH mask_default();
ALTER TABLE clinic.patients ALTER name DROP MASKED;

GRANT UNMASK ON TABLE clinic.patients TO billing_admin;               -- sees clear values
GRANT SELECT, SELECT_MASKED ON TABLE clinic.patients TO support_app;  -- may filter on masked columns, still sees masks

The other functions are mask_null (returns null), mask_hash (returns a blob hash, mainly useful in queries) and mask_replace (returns a fixed replacement). Superusers are created with UNMASK, so test masking with an ordinary role. Masking controls what a query returns; it is not encryption, and anyone with file-system access, backups or UNMASK sees clear data.

CQL and tooling additions

Smaller CQL additions remove application-side workarounds:

  • Math: abs, exp, log, log10 and round.
  • Collections: map_keys, map_values, collection_count, collection_min, collection_max, collection_sum and collection_avg, so SELECT collection_count(tags) FROM ... no longer means fetching the whole set.
  • WRITETIME and TTL on collections and UDTs now work on non-frozen collections, where each element is a separate cell; check the result shape your driver returns before relying on it.
  • Removed: the deprecated dateOf and unixTimestampOf; use toTimestamp and toUnixTimestamp. Grep your application and any CQL scripts before upgrading.

On the operations side, sstablepartitions is an offline tool for finding large partitions in SSTable files, a virtual table exposes recent log lines through CQL, the CIDR authorizer restricts roles by client network range, the crypto provider is pluggable, and more guardrails let you warn on or reject risky schema and query patterns. Check the 5.0 docs for exact option names before configuring them.

The 2038 limit and storage_compatibility_mode

Cassandra stored the local expiry time of a TTL'd cell as 32-bit seconds since 1970, which runs out in January 2038. With a 20-year maximum TTL, writes from 2018 onwards could already reach past it, and 4.x handled that with an overflow policy. 5.0's storage format moves the limit to 2106, but only once every node writes the new format, so the change is gated by storage_compatibility_mode:

  1. CASSANDRA_4 (the default after upgrade): SSTables, commit log, hints and messaging stay 4.x-compatible, 2038 is still the limit, and rollback to 4.x is possible.
  2. UPGRADING: a rolling restart with this value; once every node reports storage version 5, the 2106 limit applies.
  3. NONE: a final rolling restart; mixed-version operation is no longer possible.

Worked example: upgrading a 4.1 cluster

Worked example: a 12-node, two-datacenter 4.1 cluster on Java 11, with one legacy DTCS table, LCS on most tables and an application that calls dateOf. The plan runs in phases so each step can be checked and, until the last one, reversed.

  1. Preconditions. Confirm every node is on 4.0 or 4.1 (older versions must upgrade to 4.x first). Run a full repair. Move the DTCS table to TWCS and let it compact. Fix the application's dateOf calls. Snapshot every node and verify that you can restore one.
  2. Binary upgrade, one node at a time. nodetool drain, stop, install 5.0 with your merged cassandra.yaml (start from the 5.0 file and port your settings, rather than reusing the old file), start on Java 11 or 17, wait for UN in nodetool status and normal latency, then move on. Do not run repairs, topology changes or schema changes while versions are mixed.
  3. Soak. Run for days on 5.0 in CASSANDRA_4 mode. Rollback is still possible. Watch GC pauses, read and write latency percentiles, pending compactions and dropped messages against your 4.1 baseline.
  4. Storage version. Rolling restart with UPGRADING, then another with NONE. From here, rollback means restoring snapshots.
  5. Adopt features table by table. Switch a read-heavy table to the trie memtable and UCS L10 on one datacenter's worth of traffic, measure for a week, then widen. Select BTI as the format, and let compaction or upgradesstables rewrite files at a controlled rate. Replace legacy secondary indexes with SAI one at a time.

New clusters are simpler. 5.0 ships a cassandra_latest.yaml that sets the trie memtable, BTI with zstd compression, UCS (T4) as the default compaction, SAI as the default index, a zero-sized key cache and storage_compatibility_mode: NONE. Start new clusters from it; never drop it onto an existing cluster.

Failure modes

  • Upgrading with DTCS tables or removed functions in use: the table or the application fails after the binary swap. Both are checks you can run beforehand.
  • Flipping to NONE too early: you lose the rollback path while still finding regressions. Soak first.
  • Assuming new features are on: an upgraded cluster still runs skip-list memtables, the big format and your old compaction until you change them.
  • Mass rewrites: switching compaction or format on a large table rewrites all of its data. Throttle compaction, change one table at a time, and keep disk headroom of at least the table's size.
  • Masking mistaken for protection: superusers and roles with UNMASK see everything, and backups contain clear data.
  • Reused old yaml: settings that 5.0 renamed or removed are missed, and new options never get set.

Trade-offs

The 5.0 storage changes generally buy lower memory and GC pressure and cheaper lookups at the price of a one-way format change, and UCS buys flexibility at the price of new tuning knobs your team has not used yet. Upgrading early gets SAI, vectors and masking; waiting for a few patch releases gets other people's bug reports. A reasonable middle path is to upgrade binaries soon, stay in compatibility mode briefly, and adopt engine features one table at a time with measurements.

Further reading on this site: STCS, LCS and TWCS, the SSTable and JVM tuning on 5.0.

What to do next

  1. Inventory every table's compaction strategy and move DTCS tables to TWCS.
  2. Grep application code and scripts for dateOf and unixTimestampOf.
  3. Confirm all nodes run 4.0 or 4.1 and a supported JDK (11 or 17).
  4. Build a 5.0 cassandra.yaml from the shipped file, porting your settings deliberately.
  5. Upgrade a staging cluster, soak in CASSANDRA_4 mode, then step through UPGRADING and NONE.
  6. Pick one read-heavy table to trial the trie memtable, BTI and UCS, with before-and-after latency and GC data.
  7. Plan SAI replacements for legacy secondary indexes and enable masking on columns that support staff should not see.
Key takeaway: Cassandra 5.0 adds trie memtables, the trie-indexed BTI SSTable format, the Unified Compaction Strategy, storage-attached indexes and vectors, dynamic data masking, new CQL functions and a 2106 expiry limit, and requires Java 11 or 17. Almost all of it is opt-in on upgraded clusters. Upgrade from 4.x after removing DTCS tables and dateOf calls, soak in CASSANDRA_4 compatibility mode while rollback is still possible, step through UPGRADING and NONE, then adopt engine features one table at a time with measurements. Start new clusters from cassandra_latest.yaml.