Apache Cassandra 5.0 added a VECTOR<FLOAT, n> column type and approximate nearest neighbour (ANN) search through Storage-Attached Indexing (SAI). That lets you keep embeddings next to the rows they describe and ask for the k most similar rows in CQL, without running a separate vector database. It also brings a set of rules that surprise people the first time: your QUORUM read silently becomes ONE, paging stops working, and deleted rows can make results worse.

This article explains why. It follows an ANN query from the client to the on-disk graphs and back, using the behaviour in the Cassandra 5.0.0 source and documentation, then turns that into data-modelling advice, a sizing example, client code and an operations checklist. A general introduction to embeddings and similarity is in Cassandra vector search basics; this page goes underneath it.

Advertisement

The surface: type, index, query

Three statements cover the whole feature. A vector column holds a fixed-length array of floats; a SAI index on it makes it searchable; ORDER BY ... ANN OF ... LIMIT asks for the nearest rows.

CREATE TABLE rag.chunks (
    tenant_id   text,
    doc_id      uuid,
    chunk_no    int,
    body        text,
    lang        text,
    embedding   vector<float, 1024>,
    PRIMARY KEY ((tenant_id), doc_id, chunk_no)
);

CREATE INDEX chunks_embedding_idx ON rag.chunks (embedding) USING 'sai'
    WITH OPTIONS = {'similarity_function': 'COSINE'};
CREATE INDEX chunks_lang_idx ON rag.chunks (lang) USING 'sai';

SELECT doc_id, chunk_no, body, similarity_cosine(embedding, ?) AS score
FROM rag.chunks
WHERE tenant_id = ? AND lang = 'en'
ORDER BY embedding ANN OF ?
LIMIT 10;

The similarity function is chosen per index: COSINE (the default in 5.0.0), DOT_PRODUCT or EUCLIDEAN. The functions similarity_cosine, similarity_dot_product and similarity_euclidean compute a score in the select list so the application can apply a threshold. The documentation states that the LIMIT must be 1,000 or fewer, that results are approximate, and that you cannot ask for the least similar rows. ANN ordering only works on float vectors, and a vector index answers only ANN queries; you cannot use it in an equality predicate.

Inside one node: graphs that follow the storage engine

SAI attaches index structures to the storage engine's own units instead of keeping a separate global index. For a vector column that means one in-memory index over the memtable and one on-disk graph per SSTable segment. The graph library is JVector, a graph-based ANN engine in the DiskANN family: each vector is a node connected to a bounded number of neighbours, and a search walks from an entry point towards closer and closer nodes, keeping a beam of the best candidates seen so far.

Because the index follows the storage engine, it inherits its life cycle. A flush turns the memtable index into a new on-disk graph. Compaction merges SSTables, so it also rebuilds the vector graphs for the merged data, which is the dominant extra cost of vector indexing: building a graph means a nearest-neighbour search per inserted vector. A table with many small SSTables has many small graphs, and a query must search all of them. That is why compaction strategy matters more for vector tables than for ordinary ones. The general SAI design is covered in Storage-Attached Indexing.

Two build options are exposed in 5.0.0: maximum_node_connections (default 16, maximum 512) bounds the neighbours per node, and construction_beam_width (default 100, maximum 3200) sets how wide the search is while inserting. Larger values give better recall at the cost of slower builds and bigger graphs. A third option, optimize_for, affects queries and is covered next.

Advertisement

Following a query across the cluster

ClientANN OF q LIMIT kCoordinatorCL forced to ONE / LOCAL_ONEOne replica per token rangerange reads, each asked for top kMemtable vector indexin-memory graph, live writesSSTable segment graphsone per SSTable segmentRestricted columnsother SAI indexes, partition keyPer segment: graph walk or brute forcecandidates = multiplier x kfilterMerge, read rows, re-checkdrop deleted / overwritten entriesCoordinator mergeglobal top k by scoreNo paging: the whole top k comes back in one page.
An ANN query fans out to one replica per token range, searches every memtable and SSTable graph on each, and merges the best k at every level.

Cassandra calls these top-K queries and gives them their own validation rules in the select path. Reading the code explains each surprise.

Consistency is forced down. If the requested level needs reconciliation between replicas (QUORUM, LOCAL_QUORUM, ALL and so on), the coordinator downgrades it to ONE, or LOCAL_ONE for datacenter-local levels, and sends a client warning. SERIAL and LOCAL_SERIAL are rejected. The reason is structural: each replica returns its own top k, and two replicas that disagree about one row may each return a different set, so there is no correct way to merge them row by row. The consequence for you is that vector reads see whatever the chosen replica has, as discussed in Cassandra consistency. Write at QUORUM and keep repair running if stale results matter.

No paging. A LIMIT is mandatory, per-partition limits and aggregation are rejected, and if the driver's page size is smaller than LIMIT, the server raises the page size to the LIMIT and warns. Asking for 'the next 10' does not work; ask for a larger k once and page in the application.

Over-fetching per segment. Inside each SSTable segment the searcher asks the graph for more than k candidates, because a graph walk is approximate. The multiplier comes from optimize_for. With the default LATENCY, the multiplier is 5 at k = 1, about 1.1 at k = 100 and 1.0 at k = 1,000. With RECALL it is 10, 2.0 and 1.1. Small k therefore gets proportionally the most extra work, and switching to RECALL roughly doubles candidate work at k = 100.

Merge and re-check. Candidates from all segments and the memtable are merged by score, rows are read, and rows that were deleted or whose vector was overwritten are dropped. If many candidates are stale, fewer than k good rows survive, which is how deletes and updates degrade result quality.

Filtering and the brute-force path

Real queries combine similarity with constraints: a tenant, a language, a date range. Cassandra 5.0 accepts ANN together with restrictions only when every restricted column is indexed (or is the partition key used normally); restrictions that would need ALLOW FILTERING are rejected with 'ANN ordering by vector requires all restricted column(s) to be indexed'.

Restricting the partition key is the most effective filter, because the query touches one replica set and one partition's rows instead of every token range. For multi-tenant retrieval, putting the tenant in the partition key is the single best modelling decision you can make, as long as partitions stay a sensible size (the usual rules in Cassandra data modelling still apply).

Inside a segment, SAI also has an exact path. When the rows that pass the other restrictions, or the rows in the searched range, are few compared with the graph and the requested candidate count, the searcher skips the graph and scores those rows directly: exact nearest neighbours, no approximation. Query tracing shows which path ran, with a line reporting how many rows the search range covers and the maximum for brute force. This is good news for selective filters, and a trap for medium-selective ones: a filter matching 30 percent of a segment is too broad to brute force and can still make the graph walk discard most of the neighbours it finds.

Choosing a similarity function

Pick the function your embedding model was trained for; the model card usually says. Cosine compares directions and ignores length. Dot product equals cosine for unit-length vectors and is cheaper, so DOT_PRODUCT is a fine choice only if every vector you store and every query vector is normalised; mixing normalised and unnormalised vectors silently ranks long vectors higher. Euclidean distance is the right choice for models trained with L2 objectives. Changing the function means dropping and recreating the index, which rebuilds every graph, so decide before loading data.

Worked example: sizing 20 million chunks

A team indexes 20 million document chunks with a 1,024-dimension model, replication factor 3, in one datacenter of six nodes.

Raw vectors. 1,024 floats x 4 bytes = 4 KiB per row, 80 GB of vectors per copy and 240 GB across three replicas, roughly 40 GB per node before compression and before the text body. Compression helps little on float vectors.

Guardrails. Cassandra 5.0.0's configuration sets sai_vector_term_size_warn_threshold to 16 KiB and sai_vector_term_size_fail_threshold to 32 KiB. At 4 bytes per float, our arithmetic gives a warning above 4,096 dimensions and a failure above about 8,192 with the defaults; 1,024 is comfortably inside. Separate vector_dimensions_warn_threshold and fail guardrails exist and are disabled (-1) by default.

Graph. With 16 connections per node and 4-byte neighbour ids, adjacency alone is in the order of 64 bytes per vector, small next to the 4 KiB vector itself. The vectors, not the graph, dominate disk usage.

Query fan-out. Without a partition restriction, a query at LOCAL_ONE must touch enough replicas to cover every token range, so most nodes typically participate in each query, and each searches every SSTable graph it holds for those ranges. With the tenant in the partition key, one node answers. For a retrieval service at hundreds of queries per second, that difference decides the cluster size.

A client using the DataStax Python driver looks like this:

from cassandra.cluster import Cluster
from cassandra import ConsistencyLevel

session = Cluster(["10.0.0.11", "10.0.0.12"]).connect("rag")
ann = session.prepare(
    "SELECT doc_id, chunk_no, body, similarity_cosine(embedding, ?) AS score "
    "FROM chunks WHERE tenant_id = ? AND lang = ? "
    "ORDER BY embedding ANN OF ? LIMIT ?")
ann.consistency_level = ConsistencyLevel.LOCAL_ONE   # what the server would force anyway

def retrieve(tenant, query_vec, k=10, min_score=0.75, lang="en"):
    rows = session.execute(ann, (query_vec, tenant, lang, query_vec, k))
    hits = [r for r in rows if r.score >= min_score]
    # Fewer than k hits is normal: approximate search plus the score threshold.
    return sorted(hits, key=lambda r: r.score, reverse=True)

Setting LOCAL_ONE explicitly removes a warning per query and, more importantly, documents the real consistency in code. Measure recall offline before trusting any settings: compute exact top k with a brute-force scan over a sample, compare with the ANN result, and track recall at k as a metric, the way RAG evaluation recommends for the retrieval stage.

Data modelling patterns that hold up

  • Partition by the thing you always filter on. Tenant, project or corpus id in the partition key turns a cluster-wide search into a single-replica-set search.
  • Write vectors once. The documentation notes that ANN works best without overwriting or deleting vectors. Keep frequently updated metadata (view counts, labels) in another table keyed the same way, so updates do not leave stale graph entries.
  • Re-embed into a new column or table. A new model version means new vectors in a new index, backfilled in the background, with the application switching reads when recall is verified. Overwriting in place degrades both indexes during the transition.
  • Use TTLs carefully. Expiring rows behave like deletes: candidates found in the graph and then dropped at read time until compaction removes them.
  • Keep k small and honest. Ask for the k you will use. Large k defeats the latency multiplier and raises per-query reads on every node.

Failure modes and operations

SymptomLikely causeWhat to do
Warning: consistency downgradedQuery sent at QUORUMSet LOCAL_ONE explicitly; write at QUORUM; keep repair running
Fewer than k rows returnedStale candidates from deletes, overwrites or TTL; strict filtersSeparate mutable data; compact; raise k slightly; use RECALL
Latency grows over weeksMany SSTables, each with its own graphReview compaction; check SSTable count per table
Compaction falls behind after enabling the indexGraph building on every mergeAdd CPU headroom; schedule index creation off-peak
Recall drops after a model changeMixed vectors or wrong similarity functionSeparate column per model; verify normalisation
Query rejected, needs indexed columnsFilter on a non-indexed columnAdd a SAI index or move the column into the key

Creating a vector index on a populated table builds graphs for all existing SSTables, which is CPU-heavy; do it off-peak and watch compaction throughput. Trace a few representative queries after every schema or compaction change to see which segments were searched and whether the brute-force path ran.

When Cassandra is the right vector store

Cassandra's vector search is strongest when the embeddings belong to data that already lives in Cassandra, when queries can be scoped by partition, and when you value one system with one replication and backup story over the last few points of recall. A dedicated vector database is usually stronger when you need quantisation-heavy memory tricks over billions of vectors, rich hybrid ranking, or strongly consistent filtered search across the whole corpus. Both can be correct; the deciding question is how often your query has a natural partition to scope it.

What to do next

  1. Confirm your embedding model's intended similarity and normalisation, then choose COSINE, DOT_PRODUCT or EUCLIDEAN before loading data.
  2. Put the tenant or corpus id in the partition key and index every column you filter on with SAI.
  3. Set LOCAL_ONE explicitly on ANN statements, write at QUORUM, and keep repair healthy.
  4. Build an offline recall test with exact brute-force top k on a sample and track recall at k for each setting you change.
  5. Keep mutable metadata out of the vector table; plan re-embedding as a new column or table.
  6. Watch SSTable counts, compaction throughput and query traces after enabling the index.
  7. Only then tune optimize_for, maximum_node_connections and construction_beam_width, one at a time, against the recall test.
Key takeaway: Cassandra 5 vector search is SAI applied to embeddings: one graph per memtable and per SSTable segment, searched with over-fetching and merged at every level. That design explains the rules: top-K queries run at ONE or LOCAL_ONE, need a LIMIT, do not page, and degrade with deletes and overwrites. Scope queries by partition, write vectors once, measure recall yourself, and the feature gives you retrieval without a second database.