Cassandra is a Java process whose latency is only as good as its worst garbage-collection pause. A coordinator that stops for 800 milliseconds misses read timeouts, drops hints and looks dead to its peers, and the client sees a spike in p99 latency for every request routed through that node. For a decade JVM tuning was the most argued-about part of running Cassandra, mostly over CMS settings that no longer matter.

In 2026 the picture is simpler and different. Cassandra 5.0 runs on JDK 11 or 17 with a G1 profile that works well for most clusters, much of the memory pressure has moved off-heap, and the next major line adds JDK 21 with generational ZGC. This article explains what the defaults do, how to size memory for the whole process rather than just the heap, and how to tune from evidence in GC logs rather than folklore.

Advertisement

Where the versions stand

LineJDKsDefault collectorStatus on 2026-10-02
Cassandra 4.18, 11CMS on 8; G1 or CMS on 11 by configurationMaintained
Cassandra 5.011, 17G1 in both jvm11- and jvm17-server.optionsLatest GA (5.0.9)
Cassandra 6.0 (branch at 6.0-alpha3)11, 17, 21G1 on 11 and 17; generational ZGC on 21Pre-release; details may change

CMS was removed from the JDK in version 14, so on Cassandra 5.0 with JDK 17 the question is no longer which collector but how to configure G1, plus when it is worth evaluating ZGC. The options are split across files in conf/: jvm-server.options for flags common to every JDK, and jvm11-server.options or jvm17-server.options for version-specific flags, chosen by the start script from the running JVM.

How G1 meets a Cassandra workload

Cassandra allocates at a very high rate and most objects die young: request objects, mutations, row iterators and serialisation buffers live for one request. A small set lives long: memtable contents (when on heap), caches and schema. G1 divides the heap into equal regions, collects young regions in stop-the-world evacuation pauses, and reclaims old regions incrementally through mixed collections after a concurrent marking cycle.

The shipped G1 profile is tuned for that shape:

# conf/jvm17-server.options as shipped with Cassandra 5.0 (G1 section, active by default)
-XX:+UseG1GC
-XX:+ParallelRefProcEnabled
-XX:MaxTenuringThreshold=2
-XX:G1HeapRegionSize=16m
-XX:+UnlockExperimentalVMOptions
-XX:G1NewSizePercent=50
-XX:G1RSetUpdatingPauseTimePercent=5
-XX:MaxGCPauseMillis=300
-XX:InitiatingHeapOccupancyPercent=70
  • G1NewSizePercent=50 keeps the young generation at least half the heap so short-lived request garbage dies before promotion; it needs the experimental-options unlock on the line above.
  • MaxTenuringThreshold=2 promotes survivors quickly, because anything that survives two collections is probably a memtable object that will live until flush.
  • MaxGCPauseMillis=300 is a goal, not a guarantee. It is higher than the JVM default of 200 ms to trade some pause length for throughput.
  • InitiatingHeapOccupancyPercent=70 starts concurrent marking later than the JVM default of 45, which suits large heaps whose old generation is mostly stable.
  • G1HeapRegionSize=16m makes any single allocation of 8 MB or more a humongous object, placed directly in old regions. Large partitions and big batch mutations produce exactly these allocations.
Advertisement

Size the whole process, not the heap

If you do not set a heap, cassandra-env.sh computes one: half of system memory, capped at 15,872 MB when CMS is not in use. Set it explicitly in jvm-server.options so every node is identical and the value is reviewed. Equal -Xms and -Xmx avoid resizing.

# conf/jvm-server.options: pin the heap explicitly instead of relying on cassandra-env.sh
-Xms16G
-Xmx16G

The heap is only part of the footprint. Off-heap memory includes memtables when memtable_allocation_type is offheap_objects or offheap_buffers, the chunk cache (file_cache_size, by default the smaller of a quarter of the heap or 512 MB), networking buffers (by default the smaller of a sixteenth of the heap or 128 MB), bloom filters and compression metadata that grow with data per node, plus metaspace, thread stacks and the JIT code cache. Everything left is page cache, which is what makes SSTable reads fast.

Memory budget of a 64 GB Cassandra node (example, G1 on JDK 17)Java heap -Xmx16Gyoung (>= 50%) + old, G1 16 MB regionsOff-heap memtablesoffheap_objects: up to memtable_offheap_spaceChunk cache + networkingfile_cache_size, native buffersBloom filters, compression metadatascales with data per nodeMetaspace, thread stacks, code cacheabout 1 GBOS page cacheeverything not reserved aboveSSTable reads hit the page cache first;starving it turns reads into disk I/Oboth competeRule: heap + all off-heap + about 2 GB for the OS must fit with room for page cache, and never swap.
A worked memory budget. Heap and off-heap structures are reserved; the remainder is OS page cache, which serves SSTable reads. Over-sizing the heap shrinks the page cache.

Move pressure off the heap

The cheapest way to shorten pauses is to give the collector less to scan. The stock cassandra.yaml in 5.0 still defaults to heap_buffers, but cassandra_latest.yaml, the template for new clusters, selects offheap_objects, the trie memtable, the BTI SSTable format and the Unified Compaction Strategy. With off-heap objects the cell data of memtables lives outside the heap, so the old generation fills much more slowly between flushes, although partition and row structures stay on heap; the memtable flush article explains the flush thresholds this interacts with. The trie memtable stores memtable contents in a compact trie and is described in the configuration as significantly reducing GC load, at the cost of weaker behaviour under unevenly distributed loads.

These are table-level and node-level changes with operational consequences, so roll them out like any configuration change, one node first, watching flush frequency and off-heap usage in nodetool info.

Generational ZGC on JDK 21

ZGC does most of its work concurrently, so pause times stay in the low milliseconds regardless of heap size. On JDK 21 the generational mode has to be requested; from JDK 23 it is the default, and in JDK 24 non-generational ZGC was removed. The Cassandra 6.0 profile for JDK 21 enables it:

# cassandra-6.0 branch (6.0-alpha3): conf/jvm21-server.options, active GC lines
-XX:+UseZGC
-XX:+ZGenerational
# temporary workaround so Jamm sizes objects correctly; ZGC does not use compressed oops anyway
-XX:-UseCompressedOops

The trade-off is CPU and memory headroom. Concurrent collection uses application cores, and ZGC needs spare heap to keep allocating while it collects; when allocation outruns collection, threads stall. ZGC also does not use compressed object pointers, so the same -Xmx holds fewer objects than under G1 below 32 GB, and resident memory differs. Recalculate the budget rather than copying the G1 heap size. The background on both collectors is in the G1 and ZGC articles.

Because the JDK 21 profile ships with a release that is not yet GA, treat ZGC on Cassandra as something to benchmark on a canary, not a default to adopt everywhere.

Choosing a collector

SituationChoiceReason
Cassandra 5.0, any supported JDKG1 with the shipped profileThe only profile in the GA line; tested by the project
Strict p99 target below 50 ms, spare coresBenchmark generational ZGC once its release line is GAPauses largely independent of heap size
CPU-bound nodes near saturationG1Concurrent collection takes cores from requests
Small heap, 8 GB or lessG1ZGC headroom costs proportionally more
Still on CMS with JDK 8Upgrade path firstCMS is gone from modern JDKs; tuning it is sunk cost

Whatever the choice, the collector is rarely the root cause of bad tail latency. Large partitions, unbounded queries, tombstone-heavy reads and oversized batches create the garbage; the collector only decides how the cost is paid.

Containers and Kubernetes

Running Cassandra in containers adds one hard constraint: the kernel enforces the memory limit on the whole process, so the off-heap budget is no longer a soft concern. Size the container limit as heap plus the off-heap estimate plus JVM overhead plus margin, and request the same amount so the pod is never placed on an overcommitted node. Do not rely on container-aware percentage flags such as MaxRAMPercentage for Cassandra; an explicit -Xmx keeps the budget readable.

Page cache is charged to the container's memory cgroup too, but the kernel can reclaim it under pressure, so it does not cause kills by itself; it does mean that a tight limit silently trades away read performance. Pin CPU so concurrent GC threads and request threads are not throttled by a CPU quota, because quota throttling looks exactly like a GC pause in client latency graphs.

Read the evidence: GC logs and Cassandra signals

Tuning without measurement just moves problems around. Cassandra already reports pauses itself: the GC inspector logs any collection longer than gc_log_threshold (200 ms by default) and warns above gc_warn_threshold (1000 ms). nodetool gcstats gives the interval maximum and total; JDK Flight Recorder through jcmd <pid> JFR.start shows allocation hot spots. The unified GC log is the primary record. On JDK 11 and later, cassandra-env.sh in 5.0 already writes gc.log to the Cassandra log directory, including safepoint events, rotated across ten 10 MB files; if you put your own -Xlog:gc line in the options files, the script skips its default. A short script summarises the log:

# Summarise pauses from a unified GC log: p50, p99, max and pause time per minute
import re, statistics, sys, collections

pause = re.compile(r"\[(\d+\.\d+)s\].*Pause (Young|Remark|Cleanup|Full).*?(\d+\.\d+)ms")
pauses, per_min = [], collections.Counter()
for line in open(sys.argv[1]):
    m = pause.search(line)
    if m:
        uptime, kind, ms = float(m[1]), m[2], float(m[3])
        pauses.append((kind, ms))
        per_min[int(uptime // 60)] += ms

ms = sorted(p for _, p in pauses)
q = statistics.quantiles(ms, n=100)
print(f"pauses={len(ms)} p50={q[49]:.1f}ms p99={q[98]:.1f}ms max={ms[-1]:.1f}ms")
print("full GCs:", sum(1 for k, _ in pauses if k == "Full"))
print("worst minute: %.0f ms paused" % max(per_min.values()))

Correlate the result with nodetool tpstats dropped messages and with client p99. Pauses that coincide with large-partition reads or big batches point at the data model, not the JVM; the operational metrics article lists the counters to put beside the GC graphs.

Worked example: a 64 GB node

Consider a node with 64 GB RAM, 16 cores, NVMe disks, about 1.5 TB of data, Cassandra 5.0 on JDK 17, and p99 read latency spikes to 900 ms several times an hour. The GC log shows young pauses around 120 ms at p99 but three mixed collections per hour above 700 ms, and humongous allocation messages. The heap was 31 GB, set years ago when CMS needed it.

The fix is in three steps, each canaried on one node for a day. First, the large-partition warnings in system.log and nodetool tablehistograms show partitions over 100 MB; reading them creates humongous allocations, so the team pages the queries and plans a remodel. Second, the heap drops to 16 GB and the yaml switches to offheap_objects with explicit memtable_offheap_space; old-generation occupancy falls because memtable cells leave the heap. Third, the freed 15 GB becomes page cache, raising the cache hit ratio and cutting disk reads. Result on the canary: p99 pause 140 ms, no pause over 250 ms, and read p99 down to 35 ms. The budget now reads 16 GB heap, about 6 GB off-heap, 1 GB JVM overhead, 2 GB OS and roughly 39 GB page cache.

Failure modes and trade-offs

  • Heap too large. Above about 32 GB the JVM loses compressed oops under G1 and every reference doubles in size; big heaps also lengthen marking. More heap is rarely the answer.
  • Swap and transparent huge pages. A swapped-out heap page turns a 50 ms pause into seconds. Disable swap, and follow the production recommendations on transparent huge pages.
  • Container limits. In Kubernetes the container is killed by the kernel when heap plus off-heap exceeds the memory limit, with no Java error. Size the limit from the full budget.
  • Time to safepoint. A pause can be short while threads take long to reach the safepoint, for example during page faults on memory-mapped files. The default GC log records safepoint events, so check them before blaming the collector.
  • Tuning flags by copy. Flags copied from old blog posts, such as CMS options or a fixed young size with -Xmn, conflict with G1; the start script refuses -Xmn with G1.
  • Rolling changes too fast. A JVM change applied cluster-wide without a canary can turn a latency problem into an outage. The restart procedure in the operations article applies.

What to do next

  1. Record your Cassandra and JDK versions on every node and confirm which jvm*-server.options file each one loads.
  2. Confirm gc.log is being written on every node, keep the rotated files for at least a week, and summarise them as a baseline.
  3. Write down the full memory budget for one node: heap, every off-heap structure, JVM overhead, OS and page cache.
  4. If the heap is above 16 to 24 GB under G1, test a smaller heap together with off-heap memtables on one canary node.
  5. Find large partitions and humongous allocations before touching GC flags.
  6. When the JDK 21 profile reaches a GA release, benchmark generational ZGC on a canary with production-like load before adopting it.
Key takeaway: In 2026 Cassandra JVM tuning starts from good defaults: G1 with a large young generation on Cassandra 5.0 with JDK 11 or 17, and generational ZGC on JDK 21 in the next major line, still pre-release. Size the whole process, not just the heap, move memtables off-heap, keep the page cache large, fix large partitions before GC flags, and change one node at a time with GC logs as the evidence.