Most Cassandra GC tuning starts in the wrong place. Someone sees a p99 latency spike on a dashboard, opens jvm17-server.options, changes MaxGCPauseMillis or the heap size, restarts a node, and watches the graph for an afternoon. Often nothing improves, because the stall was never a collection, or because the collector was faithfully cleaning up garbage the application should not have produced.

This page is the diagnostic half of JVM work on Cassandra. It shows how to trace a client-side latency spike back to a specific pause on a specific node, where Cassandra allocates memory on its read, write and compaction paths, how to read G1 and safepoint log lines, which stalls look like garbage collection but are not, and how to run a tuning change as a controlled experiment rather than a guess. Collector choice, the shipped G1 flags, heap and off-heap budgets and generational ZGC are covered in Cassandra JVM tuning in 2026; read that for the configuration and this for the evidence that should drive it.

From symptom to stall

A latency spike seen by an application is the end of a chain, and each link can be checked with data you already have. Start at the client: the driver's per-node latency metrics, or a request trace, tell you whether the slow requests went through one coordinator or all of them. A cluster-wide spike points at something shared; a spike on one or two nodes points at those JVMs.

Next, on the suspect node, compare coordinator latency (ClientRequest metrics) with local read and write latency (Table metrics). If local latency is flat while coordinator latency jumps, the node was waiting on replicas or was itself frozen. Then take the timestamp of the stall and look it up in three places: the GC log, the safepoint log and the operating system's view (swap, CPU steal, disk wait). Only once you know which branch you are on does it make sense to change anything.

Client p99 spikedriver latency histogramwhich node?Coordinator latencyClientRequest metrics per nodewhich stall?Node stall windowtimestamp, durationGC pausegc.log: Pause Young / FullTime to safepointsafepoint log: ReachingOutside the JVMswap, THP, page faults, CPUcauseAllocation sourcehumongous, promotion, ratefixChange one thingcanary node, compareWalk the chain left to right; each hop narrows the search before any flag is touched.Three branches at the stall: a real collection, a slow safepoint, or the operating system.
The diagnostic chain: from the client symptom to one node, to one stall window, to one of three causes, and only then to a single controlled change.

Cassandra helps with the middle step. Its GCInspector logs every collection longer than gc_log_threshold at INFO and longer than gc_warn_threshold at WARN in system.log, so a grep for GCInspector near the spike is often the fastest first look. nodetool gcstats reports the maximum and total pause time since it was last called, which makes it useful for a quick before-and-after but useless for history. For history you need the GC log itself.

Where Cassandra allocates

Garbage collection cost is driven by two numbers: how fast the application allocates, and how much of what it allocates survives. Cassandra's paths differ sharply on both, and knowing which path produces which garbage tells you where a fix belongs.

PathWhat it allocatesLifetimeWhat makes it worse
Native transport (requests)Frames, decoded statements, result setsOne requestLarge pages, wide result rows, many small requests per second
Write pathMutations, then memtable cellsRequest, then until flushBig batches, large blobs, on-heap memtables
Read pathRow iterators, merged rows, tombstone markersOne requestWide partitions, many tombstones, ALLOW FILTERING scans
CompactionPartition iterators, buffersOne partition at a timeHigh compaction throughput, very large partitions
Repair and streamingMerkle trees, stream buffersSessionRepairing large ranges in one session
CachesKey cache, row cache entriesLongOversized row cache on heap

Two patterns cause most trouble. The first is promotion: anything that survives a couple of young collections is copied to old regions, and memtable data that lives on the heap until flush is exactly that. The more memtable data lives on heap, the more old-generation work G1 does later. Moving memtable contents off heap with memtable_allocation_type (the values include heap_buffers, offheap_buffers and offheap_objects) reduces that, at the cost of native memory that must be budgeted; see memtable flushing for how flush thresholds interact with it.

The second is humongous allocation. G1 treats any single object of at least half a region as humongous and places it directly in contiguous old regions. With the 16 MB regions Cassandra's profile uses, that is anything of 8 MB or more: a large blob in one cell, a single mutation carrying a huge batch, or a big read response buffer. Humongous objects fragment the heap and can force a full collection when no contiguous space is left. The fix is almost never a JVM flag; it is making the application stop sending objects that large.

To see the allocation profile directly rather than infer it, run an allocation profiler against a live node for a minute. With async-profiler the command is asprof -e alloc -d 60 -f alloc.html <pid>; JDK Flight Recorder through jcmd <pid> JFR.start duration=60s filename=alloc.jfr gives similar data. The flame graph shows which paths allocate the most bytes.

Reading the GC log

The unified GC log is the primary evidence. A typical young collection on JDK 17 looks like this (abridged):

[2026-10-03T10:15:02.114+0000][info][gc] GC(4521) Pause Young (Normal) (G1 Evacuation Pause) 9830M->2410M(16384M) 182.403ms
[2026-10-03T10:15:41.902+0000][info][gc] GC(4522) Pause Young (Concurrent Start) (G1 Humongous Allocation) 6120M->2988M(16384M) 211.771ms
[2026-10-03T10:15:42.010+0000][info][gc] GC(4523) Concurrent Mark Cycle
[2026-10-03T10:16:20.377+0000][info][gc] GC(4531) Pause Young (Normal) (G1 Evacuation Pause) (To-space exhausted) 15900M->15020M(16384M) 1450.214ms

Read each line as: collection number, kind, trigger, heap before and after with the total in brackets, and pause length. Pause Young (Normal) is routine. Concurrent Start means G1 has begun marking the old generation, and the trigger G1 Humongous Allocation means a large object forced it. To-space exhausted is the line to fear: G1 ran out of free regions to copy survivors into, had to handle the failure in place, and the pause grew by an order of magnitude. A Pause Full line after that means the collector fell back to compacting the whole heap. Heap after a young collection that keeps climbing between marks is old-generation growth, which on Cassandra usually means memtables, caches or a leak.

The log is easier to use when you correlate it with the latency graph. A short script turns it into a list of pauses over a threshold with their causes:

import re, sys

PAT = re.compile(r'\[(?P<ts>[^\]]+)\].*GC\((?P<n>\d+)\) (?P<kind>Pause \w+(?: \([^)]*\))*) '
                 r'(?P<before>\d+)M->(?P<after>\d+)M\((?P<total>\d+)M\) (?P<ms>[\d.]+)ms')

def pauses(path, min_ms=100.0):
    for line in open(path, encoding='utf-8', errors='replace'):
        m = PAT.search(line)
        if m and float(m['ms']) >= min_ms:
            yield m['ts'], m['kind'], int(m['before']), int(m['after']), float(m['ms'])

for ts, kind, before, after, ms in pauses(sys.argv[1]):
    flag = ' <-- to-space exhausted' if 'To-space' in kind else ''
    print(f'{ts}  {ms:8.1f} ms  {before:>6}M -> {after:>6}M  {kind}{flag}')

Run it over all rotated files, sort by time and line the output up against the spike. If the spike lines up with a long pause, you are on the GC branch. If it does not, keep looking: the safepoint log, which Cassandra's start script includes in its default GC logging on recent JDKs, records how long threads took to reach each safepoint, separately from how long the operation inside it took.

Stalls that are not garbage collection

Every stop-the-world operation, garbage collection included, runs at a safepoint: the JVM asks all application threads to stop at a known point and waits until the last one arrives. The time spent waiting is called time to safepoint, and it is not part of the GC pause figure. A safepoint log line has the shape Safepoint "G1CollectForAllocation", ... Reaching safepoint: 152340 ns, ... At safepoint: 182400000 ns, Total: .... If the reaching time is large, one thread was slow to stop, and the GC was not the problem.

On Cassandra the usual reasons are outside Java code:

  • Page faults on memory-mapped files. When SSTables are read through mmap, a thread touching a page that is not in memory blocks in the kernel until the disk returns it, and it cannot reach the safepoint while it waits. Under memory pressure with slow disks this produces stalls of hundreds of milliseconds that look like GC in client graphs. Check disk_access_mode and the page cache hit rate before blaming the collector.
  • Swap. If any part of the heap is swapped out, a collection that walks it pays disk latency per page. Cassandra's guidance is to disable swap on database hosts; confirm with free -m and vmstat 1 that si and so stay at zero.
  • Transparent huge pages. With THP set to always, the kernel can stall a thread while it compacts memory to build a huge page. Set it to madvise or never, as Cassandra's production recommendations say, and check /sys/kernel/mm/transparent_hugepage/enabled.
  • CPU starvation. On shared or throttled hosts, GC worker threads may simply not get CPU. In containers, a CPU quota smaller than the GC thread count stretches every pause; look at throttling counters in the cgroup.

The rule is simple: compare the latency spike with three numbers for the same instant, which are the pause time, the reaching-safepoint time and the operating system's wait counters. Whichever one matches the spike is the branch to work on.

Worked example: a spike every 40 seconds

A real-shaped example. A six-node cluster on Cassandra 5.0 with JDK 17 and a 16 GB G1 heap serves an order-history service. Every 40 seconds or so, p99 read latency jumps from 12 ms to over a second, but only on three nodes, and only during business hours. The keyspace uses replication factor 3.

Step one: the driver's per-host metrics confirm the two nodes. Step two: on one of them, grep GCInspector system.log shows WARN lines of 1.2 to 1.5 seconds at matching times. Step three: the pause script above lists To-space exhausted pauses, each preceded within a minute by a G1 Humongous Allocation concurrent start. So the branch is GC, and the trigger is large objects.

Step four is finding the source. An allocation profile shows the humongous allocations come from the native transport decoding large writes and from reads of large cells. The cause is a document-import job, recently moved to daytime, that stores each rendered invoice of around 10 MB as a single blob cell in a handful of customer partitions. Every write and every read of such a cell allocates buffers larger than 8 MB, and the three suffering nodes are exactly the replicas of those partitions.

Step five is the fix, and it is not a JVM flag. The batch-size thresholds do not help here, because this is one large cell rather than a multi-partition batch. The job is changed to put documents in object storage and keep only a pointer and metadata in Cassandra, and the humongous triggers disappear from the log. Raising the heap or the region size would have hidden the symptom for a while and made each pause longer when it came back. Oversized multi-partition batches cause the same pattern; Cassandra batches explains when batches are correct.

Running a tuning change as an experiment

When the evidence does point at configuration, change it the way you would ship code: one variable at a time, on one node, against a baseline. A tuning experiment has four parts.

  1. Baseline. Record a day of GC log, p99 and p999 local latency, and the allocation rate from the log (heap before a young collection minus heap after the previous one, divided by the interval). Note the workload mix so you can compare like with like.
  2. Canary. Apply the change to one node through the options file, restart it, and leave its peers unchanged. Because all replicas see similar traffic, the peers are your control group.
  3. Load. Let production traffic run for a full daily cycle, or replay a representative workload with cassandra-stress in a staging cluster when the change is risky. Include compaction and repair, which change the allocation profile.
  4. Compare. Put the canary's pause distribution and local latency beside two peers. Judge the tail, not the average.

Keep the change only if the canary wins on the metric you care about without losing elsewhere; a lower pause target that doubles GC CPU time can make throughput worse. The metrics to graph during the experiment are described in Cassandra operational metrics.

Failure modes and trade-offs

Failure modes worth knowing before they happen:

  • Bigger heap, longer pauses. Raising the heap to make pauses rarer also gives each mixed or full collection more to do. If pauses are caused by an allocation pattern, a larger heap delays and lengthens them.
  • Pause target too low. Setting MaxGCPauseMillis very low makes G1 shrink the young generation, which increases collection frequency and promotion, and can end in more to-space exhaustion, not less.
  • Native memory forgotten. Off-heap memtables, the chunk cache, bloom filters and direct buffers all live outside the heap. Moving work off heap without budgeting for it trades GC pauses for the kernel's OOM killer.
  • Masking the cause. Large partitions, tombstone-heavy reads and big batches are data-model problems. GC tuning can soften them but not remove them; tombstones in particular inflate read-path garbage.

The trade-off running through all of these is between pause length, pause frequency and CPU spent collecting. You can move cost between the three, but the only way to remove it is to allocate less.

What to do next

A checklist you can work through on your own cluster this week:

  • Confirm GC logging, including safepoint events, is on for every node and that the logs rotate rather than overwrite within a day.
  • Run the pause script over a day of logs and list pauses above your latency SLO with their triggers.
  • Grep system.log for GCInspector WARN lines; note which nodes and times they cluster on.
  • Check swap, transparent huge pages and disk_access_mode on every host before touching JVM flags.
  • Take a 60-second allocation profile during peak load and identify the top three allocating paths.
  • Fix application-side causes first: large blobs, oversized batches, wide partitions and tombstone-heavy reads.
  • If a flag change is still warranted, run it as a one-node canary for a full daily cycle and compare tail pauses against two peers.
Key takeaway: Treat a Cassandra latency spike as a chain to walk, not a flag to change. Find the node, find the stall window, and decide whether it was a real collection, a slow safepoint or the operating system. When it is GC, find the allocation that caused it, usually large blobs, oversized batches, wide partitions or on-heap memtables, and fix that first. Change JVM flags only as one-node canary experiments judged on tail pauses against unchanged peers.