Every Cassandra SSTable carries a bloom filter over the partition keys it contains. On a read for one partition, the node asks each candidate SSTable's filter whether the key might be there; a 'no' skips the file without touching its index or data, and a 'maybe' sends the read on to the index. The filter can say 'maybe' for a key that is absent, never 'no' for a key that is present. One table property, bloom_filter_fp_chance, sets how often that false 'maybe' happens, and it decides how much off-heap memory the filters take.
What a bloom filter is and how it is laid out inside an SSTable is covered in Cassandra bloom filters. This page is about operating them: how to size them, how to pick the false-positive chance for a given compaction strategy and read pattern, how to change it safely on a live cluster and how to read the numbers Cassandra reports. Defaults and commands below were checked against the Apache Cassandra documentation on 2026-10-02.
Where the filter sits and what a false positive costs
A single-partition read gathers data from the memtable and from every SSTable whose token range covers the key. For each of those SSTables, Cassandra checks the filter first. The filter is keyed on the partition key only, so it answers 'could this file contain partition K', not 'could it contain row R of K'. Clustering columns, cell values and secondary index terms are invisible to it.
A true 'no' saves the most expensive part of the read for that file: the partition index lookup and the seek into the data component. A false positive does exactly that work and finds nothing. On a node whose index and data are hot in the page cache the waste is CPU and a little latency; on a node that must go to disk it can be a real I/O per false positive, which is why the effect shows up first in tail latency.
Two kinds of reads never benefit. Range scans and token-range queries, including full-table scans by Spark or a bulk exporter, walk SSTables by token order, and a filter of individual keys cannot rule out a range. And reads of keys that really exist in many SSTables, because the partition has been updated across many flushes, pay for every file that holds a piece of it; no filter setting fixes that, compaction does. The read path as a whole, including the caches and indexes after the filter, is explained in the Cassandra read path in depth.
The sizing math
For a filter built for n keys with false-positive probability p, the optimal number of bits is about n times -ln(p) / (ln 2)^2, and the optimal number of hash functions is the bits per key times ln 2. The memory therefore grows with the logarithm of 1/p: each tenfold reduction in p costs roughly 4.8 more bits per key. Two facts make this matter more than it looks.
First, n is the number of partition keys in each SSTable, summed over all SSTables, not the number of distinct partitions on the node. A partition present in four files appears in four filters. Second, the filter size depends on the key count, not the data volume: a table of a billion tiny partitions needs far bigger filters than a table of a million huge ones holding the same bytes. Partition key design therefore drives filter memory, which is one more reason the trade-offs in compaction strategy choice and data modelling interact.
The documentation's own figure is that a filter at 0.01 needs about three times the memory of one at 0.1. The textbook formula gives about twice, because 9.6 bits per key against 4.8. Treat the formula as a planning estimate and the measured Bloom filter space used as the truth; the implementation's choices, not the ideal formula, decide the bytes you actually pay for.
import math
def bits_per_key(p):
"""Optimal bits per key for false-positive probability p (textbook formula)."""
return -math.log(p) / (math.log(2) ** 2)
def hashes(p):
return bits_per_key(p) * math.log(2)
def filter_gib(keys_in_sstables, p):
"""Keys counted once per SSTable that holds them, not once per node."""
return keys_in_sstables * bits_per_key(p) / 8 / 2**30
def wasted_probes(candidate_sstables, sstables_holding_key, p):
return (candidate_sstables - sstables_holding_key) * p
for p in (0.1, 0.01, 0.001):
print(f"p={p:<6} bits/key={bits_per_key(p):5.1f} k={hashes(p):4.1f} "
f"1.5e9 keys={filter_gib(1.5e9, p):4.2f} GiB "
f"STCS(30 files)={wasted_probes(30, 2, p):.2f} "
f"TWCS(200 files)={wasted_probes(200, 1, p):.2f}")Running it prints:
p=0.1 bits/key= 4.8 k= 3.3 1.5e9 keys=0.84 GiB STCS(30 files)=2.80 TWCS(200 files)=19.90
p=0.01 bits/key= 9.6 k= 6.6 1.5e9 keys=1.67 GiB STCS(30 files)=0.28 TWCS(200 files)=1.99
p=0.001 bits/key= 14.4 k=10.0 1.5e9 keys=2.51 GiB STCS(30 files)=0.03 TWCS(200 files)=0.20The last two columns are wasted index lookups per read for a key in 2 of 30 STCS files or 1 of 200 TWCS windows; the memory column is their price.
Choosing fp_chance by compaction strategy and read pattern
Cassandra's default is 0.1 for tables using LeveledCompactionStrategy and 0.01 for everything else. The reason for the difference follows from the wasted-probe formula. Under LCS, SSTables in each level above L0 have non-overlapping token ranges, so a read for one key has at most one candidate file per level plus whatever sits in L0. With a handful of levels, even a filter wrong one time in ten wastes well under one lookup per read, and LCS tables are often read-heavy with many keys, where saving memory matters. Under size-tiered and time-window strategies, every SSTable typically covers the whole token range of the node, so every file is a candidate for every key and the false-positive rate is multiplied by the file count.
| Workload | Candidate files per read | Starting point | Why |
|---|---|---|---|
| LCS, point reads | about one per level, plus L0 | 0.1 (default) | Few candidates; memory saving outweighs rare wasted lookups |
| STCS or UCS, point reads | tens of files | 0.01 (default) | Rate is multiplied by file count |
| TWCS, point reads across all history | one per time window, often hundreds | 0.01, consider 0.001 | Hundreds of candidates make 0.01 cost about two wasted lookups per read |
| Write-mostly, read by scans only | irrelevant: scans do not use the filter | 0.1 or higher | Memory spent on filters buys nothing |
| Point reads, memory-starved node | any | measure first | Raising p may be cheaper than adding RAM, but check tail latency |
Two notes on the table. Unified compaction in Cassandra 5.0 is not LCS, so by the documentation's 'all other cases' wording it should get 0.01; confirm with DESCRIBE TABLE. And the TWCS line assumes reads that genuinely search all of history for one key; if your reads restrict clustering ranges to recent time, Cassandra can sometimes skip old SSTables using their stored metadata before the filter is consulted, so measure SSTables per read before deciding the filter is the problem.
Memory: off-heap, but not free
Bloom filters live in RAM, but off the Java heap, so the documentation tells operators not to count them when choosing the maximum heap size. They still count everywhere else: in the process's resident memory, against a container memory limit and against the page cache that serves your data reads. A node with a 16 GB heap, 3 GB of bloom filters and other off-heap structures in a 32 GB container has much less page cache than its owner thinks.
Budget filters per node as the sum over tables of Bloom filter off heap memory used, and track it as data grows. Filter memory rises with every new SSTable until compaction merges duplicate keys away, so a node that falls behind on compaction holds more filter memory than its data alone would suggest. Sizing the whole process, heap and off-heap together, is covered in Cassandra JVM tuning in 2026.
Changing the setting on a live cluster
The filter is computed when an SSTable is written and persisted as the SSTable's Filter component. ALTER TABLE changes the setting for files written from then on; existing SSTables keep their old filters until compaction rewrites them. On a table that compacts actively, the change spreads within days. On a large STCS or TWCS table whose biggest files may not be compacted for weeks, the old filters can stay for a long time, so force the rewrite.
-- 1. Change the setting. Only SSTables written from now on use it.
ALTER TABLE audit.events WITH bloom_filter_fp_chance = 0.1;
# 2. Rewrite existing SSTables on one node at a time, so their filters are rebuilt.
# -a includes SSTables already on the current format; -j limits parallel rewrites.
nodetool upgradesstables -a -j 2 audit events
# 3. Confirm the memory moved before going to the next node.
nodetool tablestats audit.events | grep -i "bloom filter"The rewrite is a compaction of every SSTable in the table: it reads and writes all of the table's data on that node, competes with live compaction for I/O and needs free disk space for the files being rewritten. Do it one node at a time, ideally one rack at a time, with a low -j value, and watch read latency on the node while it runs. Because the property is schema, the ALTER applies cluster-wide at once; only the rewrite is staged. Lowering p increases memory the moment new files appear, so check headroom before lowering it on a big table.
Reading what Cassandra reports
nodetool tablestats keyspace.table prints four bloom filter lines per table, and the same values are exposed as table metrics over JMX. Sorting by them with nodetool tablestats -s bloom_filter_off_heap_memory_used lists the tables using the most filter memory.
Bloom filter false positives: 18342
Bloom filter false ratio: 0.00791
Bloom filter space used: 1803442112
Bloom filter off heap memory used: 1803440960- False positives is a counter of filter 'maybe' answers that turned out wrong. Watch its rate, not its absolute value.
- False ratio relates false positives to all positive answers the filters gave, so it is not the same quantity as
bloom_filter_fp_chance. A table whose reads usually find the key in the first file checked will show a low ratio; one whose reads hit many files that lack the key will show a ratio close to the configured chance. Compare it with the setting as a sanity check, not as a target. - Space used and off heap memory used are the size of the Filter components on disk and in memory; they should track each other.
Pair these with nodetool tablehistograms keyspace.table, whose SSTables column shows how many files reads touched at each percentile. A high SSTables-per-read count with a low false-positive rate means the key really is spread over many files, and the fix is compaction or data modelling, not the filter.
Worked example 1: freeing memory on an append-only table
An audit table on STCS receives 40,000 inserts a second and is read only by a nightly Spark export that scans token ranges, plus an occasional support lookup by event ID. Each node holds about 1.5 billion partition keys across its SSTables. tablestats shows about 1.8 GB of off-heap filter memory at the default 0.01, close to what the calculator predicts. The nodes are tight on page cache and the hot tables are suffering.
The scans gain nothing from filters, and support lookups are rare enough that a few wasted index probes do not matter. The team sets bloom_filter_fp_chance to 0.1, runs upgradesstables -a -j 2 on one node, sees filter memory fall to roughly a third, consistent with the documentation's ratio, and the false-positive rate rise with no measurable change in the support lookup latency. They roll the rewrite through the remaining nodes one per night.
Worked example 2: tightening a hot lookup table
A device-state table on TWCS with daily windows and 180 days of retention keeps about 180 SSTables per node. The API reads the latest state of a device by partition key, and the device may have last reported months ago, so reads cannot be restricted to recent windows. p99 latency is high and the false-positive rate is large. With 180 candidates and p = 0.01, each read wastes about 1.8 index lookups; at 0.001 it wastes about 0.18. The table has 400 million keys per node, so the extra memory at 0.001 is a few hundred megabytes, which the node can afford.
They lower the setting and rewrite old windows with upgradesstables during quiet hours. The better long-term fix is the data model: a latest-state table keyed by device makes each read hit one or two files.
Failure modes
- ALTER without a rewrite: the setting changes but old SSTables keep their filters, and nobody notices for weeks.
- Rewrite everywhere at once: every node compacting all data at the same time raises read latency cluster-wide and can fill disks.
- Lowering p on a huge table without headroom: off-heap memory rises with each new file until the container is killed for exceeding its limit.
- Tuning for scans: spending memory on filters for tables read only by range or token scans.
What to do next
- Run
nodetool tablestats -s bloom_filter_off_heap_memory_usedon one node and list the five tables using the most filter memory. - For each, write down the compaction strategy, how it is read (point reads or scans) and the SSTables-per-read percentiles from
tablehistograms. - Raise
bloom_filter_fp_chanceon tables read only by scans; consider lowering it on TWCS or STCS tables with many candidate files and point reads. - After each
ALTER, rewrite withupgradesstables -aone node at a time, and confirm the space used changed before moving on. - Add filter off-heap memory to your per-node memory budget alongside heap and page cache, and alert on its growth.