Apache Cassandra 5.0 added a VECTOR<FLOAT, n> column type and approximate nearest neighbour (ANN) search through Storage-Attached Indexing (SAI). That lets you keep embeddings next to the rows they describe and ask for the k most similar rows in CQL, without running a separate vector database. It also brings a set of rules that surprise people the first time: your QUORUM read silently becomes ONE, paging stops working, and deleted rows can make results worse.
This article explains why. It follows an ANN query from the client to the on-disk graphs and back, using the behaviour in the Cassandra 5.0.0 source and documentation, then turns that into data-modelling advice, a sizing example, client code and an operations checklist. A general introduction to embeddings and similarity is in Cassandra vector search basics; this page goes underneath it.
The surface: type, index, query
Three statements cover the whole feature. A vector column holds a fixed-length array of floats; a SAI index on it makes it searchable; ORDER BY ... ANN OF ... LIMIT asks for the nearest rows.
CREATE TABLE rag.chunks (
tenant_id text,
doc_id uuid,
chunk_no int,
body text,
lang text,
embedding vector<float, 1024>,
PRIMARY KEY ((tenant_id), doc_id, chunk_no)
);
CREATE INDEX chunks_embedding_idx ON rag.chunks (embedding) USING 'sai'
WITH OPTIONS = {'similarity_function': 'COSINE'};
CREATE INDEX chunks_lang_idx ON rag.chunks (lang) USING 'sai';
SELECT doc_id, chunk_no, body, similarity_cosine(embedding, ?) AS score
FROM rag.chunks
WHERE tenant_id = ? AND lang = 'en'
ORDER BY embedding ANN OF ?
LIMIT 10;The similarity function is chosen per index: COSINE (the default in 5.0.0), DOT_PRODUCT or EUCLIDEAN. The functions similarity_cosine, similarity_dot_product and similarity_euclidean compute a score in the select list so the application can apply a threshold. The documentation states that the LIMIT must be 1,000 or fewer, that results are approximate, and that you cannot ask for the least similar rows. ANN ordering only works on float vectors, and a vector index answers only ANN queries; you cannot use it in an equality predicate.
Inside one node: graphs that follow the storage engine
SAI attaches index structures to the storage engine's own units instead of keeping a separate global index. For a vector column that means one in-memory index over the memtable and one on-disk graph per SSTable segment. The graph library is JVector, a graph-based ANN engine in the DiskANN family: each vector is a node connected to a bounded number of neighbours, and a search walks from an entry point towards closer and closer nodes, keeping a beam of the best candidates seen so far.
Because the index follows the storage engine, it inherits its life cycle. A flush turns the memtable index into a new on-disk graph. Compaction merges SSTables, so it also rebuilds the vector graphs for the merged data, which is the dominant extra cost of vector indexing: building a graph means a nearest-neighbour search per inserted vector. A table with many small SSTables has many small graphs, and a query must search all of them. That is why compaction strategy matters more for vector tables than for ordinary ones. The general SAI design is covered in Storage-Attached Indexing.
Two build options are exposed in 5.0.0: maximum_node_connections (default 16, maximum 512) bounds the neighbours per node, and construction_beam_width (default 100, maximum 3200) sets how wide the search is while inserting. Larger values give better recall at the cost of slower builds and bigger graphs. A third option, optimize_for, affects queries and is covered next.
Following a query across the cluster
Cassandra calls these top-K queries and gives them their own validation rules in the select path. Reading the code explains each surprise.
Consistency is forced down. If the requested level needs reconciliation between replicas (QUORUM, LOCAL_QUORUM, ALL and so on), the coordinator downgrades it to ONE, or LOCAL_ONE for datacenter-local levels, and sends a client warning. SERIAL and LOCAL_SERIAL are rejected. The reason is structural: each replica returns its own top k, and two replicas that disagree about one row may each return a different set, so there is no correct way to merge them row by row. The consequence for you is that vector reads see whatever the chosen replica has, as discussed in Cassandra consistency. Write at QUORUM and keep repair running if stale results matter.
No paging. A LIMIT is mandatory, per-partition limits and aggregation are rejected, and if the driver's page size is smaller than LIMIT, the server raises the page size to the LIMIT and warns. Asking for 'the next 10' does not work; ask for a larger k once and page in the application.
Over-fetching per segment. Inside each SSTable segment the searcher asks the graph for more than k candidates, because a graph walk is approximate. The multiplier comes from optimize_for. With the default LATENCY, the multiplier is 5 at k = 1, about 1.1 at k = 100 and 1.0 at k = 1,000. With RECALL it is 10, 2.0 and 1.1. Small k therefore gets proportionally the most extra work, and switching to RECALL roughly doubles candidate work at k = 100.
Merge and re-check. Candidates from all segments and the memtable are merged by score, rows are read, and rows that were deleted or whose vector was overwritten are dropped. If many candidates are stale, fewer than k good rows survive, which is how deletes and updates degrade result quality.
Filtering and the brute-force path
Real queries combine similarity with constraints: a tenant, a language, a date range. Cassandra 5.0 accepts ANN together with restrictions only when every restricted column is indexed (or is the partition key used normally); restrictions that would need ALLOW FILTERING are rejected with 'ANN ordering by vector requires all restricted column(s) to be indexed'.
Restricting the partition key is the most effective filter, because the query touches one replica set and one partition's rows instead of every token range. For multi-tenant retrieval, putting the tenant in the partition key is the single best modelling decision you can make, as long as partitions stay a sensible size (the usual rules in Cassandra data modelling still apply).
Inside a segment, SAI also has an exact path. When the rows that pass the other restrictions, or the rows in the searched range, are few compared with the graph and the requested candidate count, the searcher skips the graph and scores those rows directly: exact nearest neighbours, no approximation. Query tracing shows which path ran, with a line reporting how many rows the search range covers and the maximum for brute force. This is good news for selective filters, and a trap for medium-selective ones: a filter matching 30 percent of a segment is too broad to brute force and can still make the graph walk discard most of the neighbours it finds.
Choosing a similarity function
Pick the function your embedding model was trained for; the model card usually says. Cosine compares directions and ignores length. Dot product equals cosine for unit-length vectors and is cheaper, so DOT_PRODUCT is a fine choice only if every vector you store and every query vector is normalised; mixing normalised and unnormalised vectors silently ranks long vectors higher. Euclidean distance is the right choice for models trained with L2 objectives. Changing the function means dropping and recreating the index, which rebuilds every graph, so decide before loading data.
Worked example: sizing 20 million chunks
A team indexes 20 million document chunks with a 1,024-dimension model, replication factor 3, in one datacenter of six nodes.
Raw vectors. 1,024 floats x 4 bytes = 4 KiB per row, 80 GB of vectors per copy and 240 GB across three replicas, roughly 40 GB per node before compression and before the text body. Compression helps little on float vectors.
Guardrails. Cassandra 5.0.0's configuration sets sai_vector_term_size_warn_threshold to 16 KiB and sai_vector_term_size_fail_threshold to 32 KiB. At 4 bytes per float, our arithmetic gives a warning above 4,096 dimensions and a failure above about 8,192 with the defaults; 1,024 is comfortably inside. Separate vector_dimensions_warn_threshold and fail guardrails exist and are disabled (-1) by default.
Graph. With 16 connections per node and 4-byte neighbour ids, adjacency alone is in the order of 64 bytes per vector, small next to the 4 KiB vector itself. The vectors, not the graph, dominate disk usage.
Query fan-out. Without a partition restriction, a query at LOCAL_ONE must touch enough replicas to cover every token range, so most nodes typically participate in each query, and each searches every SSTable graph it holds for those ranges. With the tenant in the partition key, one node answers. For a retrieval service at hundreds of queries per second, that difference decides the cluster size.
A client using the DataStax Python driver looks like this:
from cassandra.cluster import Cluster
from cassandra import ConsistencyLevel
session = Cluster(["10.0.0.11", "10.0.0.12"]).connect("rag")
ann = session.prepare(
"SELECT doc_id, chunk_no, body, similarity_cosine(embedding, ?) AS score "
"FROM chunks WHERE tenant_id = ? AND lang = ? "
"ORDER BY embedding ANN OF ? LIMIT ?")
ann.consistency_level = ConsistencyLevel.LOCAL_ONE # what the server would force anyway
def retrieve(tenant, query_vec, k=10, min_score=0.75, lang="en"):
rows = session.execute(ann, (query_vec, tenant, lang, query_vec, k))
hits = [r for r in rows if r.score >= min_score]
# Fewer than k hits is normal: approximate search plus the score threshold.
return sorted(hits, key=lambda r: r.score, reverse=True)Setting LOCAL_ONE explicitly removes a warning per query and, more importantly, documents the real consistency in code. Measure recall offline before trusting any settings: compute exact top k with a brute-force scan over a sample, compare with the ANN result, and track recall at k as a metric, the way RAG evaluation recommends for the retrieval stage.
Data modelling patterns that hold up
- Partition by the thing you always filter on. Tenant, project or corpus id in the partition key turns a cluster-wide search into a single-replica-set search.
- Write vectors once. The documentation notes that ANN works best without overwriting or deleting vectors. Keep frequently updated metadata (view counts, labels) in another table keyed the same way, so updates do not leave stale graph entries.
- Re-embed into a new column or table. A new model version means new vectors in a new index, backfilled in the background, with the application switching reads when recall is verified. Overwriting in place degrades both indexes during the transition.
- Use TTLs carefully. Expiring rows behave like deletes: candidates found in the graph and then dropped at read time until compaction removes them.
- Keep k small and honest. Ask for the k you will use. Large k defeats the latency multiplier and raises per-query reads on every node.
Failure modes and operations
| Symptom | Likely cause | What to do |
|---|---|---|
| Warning: consistency downgraded | Query sent at QUORUM | Set LOCAL_ONE explicitly; write at QUORUM; keep repair running |
| Fewer than k rows returned | Stale candidates from deletes, overwrites or TTL; strict filters | Separate mutable data; compact; raise k slightly; use RECALL |
| Latency grows over weeks | Many SSTables, each with its own graph | Review compaction; check SSTable count per table |
| Compaction falls behind after enabling the index | Graph building on every merge | Add CPU headroom; schedule index creation off-peak |
| Recall drops after a model change | Mixed vectors or wrong similarity function | Separate column per model; verify normalisation |
| Query rejected, needs indexed columns | Filter on a non-indexed column | Add a SAI index or move the column into the key |
Creating a vector index on a populated table builds graphs for all existing SSTables, which is CPU-heavy; do it off-peak and watch compaction throughput. Trace a few representative queries after every schema or compaction change to see which segments were searched and whether the brute-force path ran.
When Cassandra is the right vector store
Cassandra's vector search is strongest when the embeddings belong to data that already lives in Cassandra, when queries can be scoped by partition, and when you value one system with one replication and backup story over the last few points of recall. A dedicated vector database is usually stronger when you need quantisation-heavy memory tricks over billions of vectors, rich hybrid ranking, or strongly consistent filtered search across the whole corpus. Both can be correct; the deciding question is how often your query has a natural partition to scope it.
What to do next
- Confirm your embedding model's intended similarity and normalisation, then choose COSINE, DOT_PRODUCT or EUCLIDEAN before loading data.
- Put the tenant or corpus id in the partition key and index every column you filter on with SAI.
- Set LOCAL_ONE explicitly on ANN statements, write at QUORUM, and keep repair healthy.
- Build an offline recall test with exact brute-force top k on a sample and track recall at k for each setting you change.
- Keep mutable metadata out of the vector table; plan re-embedding as a new column or table.
- Watch SSTable counts, compaction throughput and query traces after enabling the index.
- Only then tune optimize_for, maximum_node_connections and construction_beam_width, one at a time, against the recall test.