Heap tuning has a reputation as a dark art of dozens of flags. In practice, on a modern JDK with G1, almost all of the benefit comes from three decisions made from measurements: how big the heap should be, what pause time to ask for, and what to do when memory runs out. Most of the remaining flags exist for edge cases, and setting them without evidence usually makes things worse, because they override the adaptive sizing that G1 does on its own.
This article is a procedure rather than a flag list. It shows how to measure the live set and the allocation rate, turns those numbers into a heap size with a worked example, explains the few G1 settings worth touching and why, and covers humongous objects, OutOfMemoryError handling and the failures you will meet. The memory layout itself, compressed pointers and why process memory exceeds -Xmx are covered in JVM memory in depth, and container memory budgets in Java in containers.
What you are actually tuning
The heap holds every Java object. A garbage collector reclaims space from objects that are no longer reachable, and the amount of free space determines how often it must run. Two quantities describe the workload. The live set is the memory occupied by objects that are still reachable at steady state, such as caches, session state and in-flight requests. The allocation rate is how fast new objects are created, usually hundreds of megabytes per second for a busy service, most of which die young.
Heap size trades memory for GC work. A heap barely larger than the live set forces constant collection and, eventually, OutOfMemoryError. A much larger heap collects less often, but costs memory and can lengthen some pauses and the time to fault pages in. Your job is to find the size where GC work and pauses meet your goals with headroom for spikes, then stop.
Step 1: turn on the evidence
Every tuning decision should come from a GC log captured under realistic peak load, not from a laptop. Unified logging is cheap enough to leave on in production with rotation, and a heap dump on OutOfMemoryError is the only reliable way to explain one after the fact. Make sure the dump directory has free space at least as large as the maximum heap.
# Unified GC logging (JDK 9+): rotate 5 files of 20 MB each.
java -Xlog:gc*,gc+heap=debug:file=/var/log/app/gc.log:time,uptime,level,tags:filecount=5,filesize=20m \
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps \
-jar app.jar
# On a running JVM, without restarting it:
jcmd <pid> GC.heap_info # current occupancy by region type
jcmd <pid> GC.class_histogram # live objects by class (triggers a full GC: not at peak!)
jcmd <pid> VM.flags -all | grep -E "MaxHeapSize|InitialHeapSize|MaxGCPauseMillis|G1HeapRegionSize"Note that a class histogram forces a full collection to count only live objects, so take it off-peak or on a canary. GC.heap_info is safe at any time.
Step 2: measure the live set and allocation rate
In a G1 log, each pause line records heap occupancy before and after the pause. After young collections the number still includes old regions full of dead objects that have not yet been collected, so it overstates the live set. After mixed collections, which also reclaim old regions, and after any full collection, the number is close to the real live set. Take the highest such value at peak load as your live-set estimate. This short script extracts it, together with pause percentiles and the interval between pauses:
import re, statistics, sys
# Example G1 lines:
# [2026-10-02T10:00:01.123+0000][12.345s][info][gc] GC(41) Pause Young (Normal) (G1 Evacuation Pause) 2900M->1350M(4096M) 18.734ms
# [2026-10-02T10:00:09.001+0000][20.223s][info][gc] GC(44) Pause Young (Mixed) (G1 Evacuation Pause) 2100M->1180M(4096M) 24.102ms
PAUSE = re.compile(r"\[(\d+\.\d+)s\].*Pause (Young|Full)[^)]*\)(?: \([^)]*\))* (\d+)M->(\d+)M\((\d+)M\) ([\d.]+)ms")
events = []
for line in open(sys.argv[1], encoding="utf-8"):
m = PAUSE.search(line)
if m:
t, kind, before, after, cap, ms = m.groups()
events.append((float(t), kind, int(before), int(after), float(ms), "(Mixed)" in line)) # not "(Prepare Mixed)"
after_mixed = [e[3] for e in events if e[5] or e[1] == "Full"]
pauses = sorted(e[4] for e in events)
gaps = [b[0] - a[0] for a, b in zip(events, events[1:])]
print("live set estimate (MiB):", max(after_mixed) if after_mixed else "need mixed or full GCs")
print("p50 / p99 pause (ms):", pauses[len(pauses) // 2], pauses[int(len(pauses) * 0.99)])
print("median seconds between pauses:", statistics.median(gaps))The allocation rate is the growth between pauses: occupancy before one young pause minus occupancy after the previous one, divided by the gap. If the previous pause left 1,350 MiB and the next starts at 2,900 MiB 4 seconds later, 1,550 MiB was allocated in 4 seconds, about 390 MiB per second. Freed space per pause is not the same thing, because promoted objects are not freed. Average over many pauses and look at peak hours; allocation rate tracks traffic.
Step 3: size the heap
A common starting heuristic for G1, not a JDK rule, is a maximum heap of three to four times the live set. That leaves room for the young generation, for garbage accumulating in old regions between concurrent cycles, and for spikes. For a measured live set of 1.2 GiB, three times gives 3.6 GiB and four times gives 4.8 GiB, so 4 GiB is a reasonable first choice. Validate it under load and adjust: if old-generation occupancy after mixed collections stays well under half the heap and pauses meet the goal, you may be able to shrink it; if concurrent cycles run back to back, grow it.
Then check the young-collection frequency. G1 sizes the young generation adaptively, by default between 5 and 60 percent of the heap, to meet the pause goal. Suppose it settles on 1.5 GiB of eden. At 390 MiB per second of allocation, eden fills in 1,536 / 390, about 3.9 seconds, so you will see a young pause roughly every four seconds. That is healthy. A pause several times per second suggests either a very high allocation rate, which a profiler such as async-profiler in allocation mode will explain, or a pause goal so tight that G1 shrinks eden to almost nothing.
For servers, set -Xms equal to -Xmx. The heap then never resizes, the footprint is predictable for capacity planning and container limits, and the GC does not spend cycles growing and shrinking. Add -XX:+AlwaysPreTouch when startup time is cheaper than latency: it touches every heap page at startup so the first minutes of traffic do not pay page faults. It makes startup slower on large heaps.
Step 4: set the pause goal, and leave the young generation alone
G1 takes a pause-time goal with -XX:MaxGCPauseMillis, 200 ms by default. It is a goal, not a guarantee: G1 chooses the young generation size and how many old regions to add to each mixed collection so that predicted pauses stay under it. Lowering it to 100 or 50 ms gives shorter but more frequent pauses and slightly lower throughput. Raising it does the opposite. Pick it from your latency budget, then verify the p99 pause in the log.
Do not set -Xmn or -XX:NewRatio with G1. Fixing the young generation size takes away the main lever G1 uses to meet the pause goal, and the pause goal is then effectively ignored. This is one of the most common harmful settings copied from old CMS-era tuning guides. G1 in depth explains the regions, remembered sets and pause anatomy behind this.
Put together, a starting configuration for the 1.2 GiB live-set service from step 3 looks like this:
# A starting point for a latency-sensitive service with a measured live set of about 1.2 GiB.
# -Xms = -Xmx: no resizing. MaxGCPauseMillis: the goal G1 sizes the young generation to meet.
# AlwaysPreTouch: fault in heap pages at startup, not during traffic.
# ExitOnOutOfMemoryError: let the orchestrator restart a broken JVM (the dump is written first).
java -XX:+UseG1GC -Xms4g -Xmx4g \
-XX:MaxGCPauseMillis=100 \
-XX:+AlwaysPreTouch \
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps \
-XX:+ExitOnOutOfMemoryError \
-Xlog:gc*:file=/var/log/app/gc.log:time,uptime,level,tags:filecount=5,filesize=20m \
-jar app.jar
# Deliberately absent: -Xmn and NewRatio (they fix the young size and disable G1's pause adaptation),
# and IHOP overrides (adaptive IHOP is on by default; change it only with log evidence).
Step 5: the old generation, IHOP and evacuation failure
G1 starts a concurrent marking cycle when old-generation occupancy crosses the initiating heap occupancy threshold. The starting value of -XX:InitiatingHeapOccupancyPercent is 45, and adaptive IHOP, on by default, adjusts the real threshold from observed allocation and marking times. Marking finds dead old objects, and the following mixed collections reclaim them. If marking starts too late, or the heap is simply too small, G1 runs out of free regions to copy surviving objects into. The log reports this as an evacuation failure or to-space exhausted, and it is usually followed by a slow full collection.
When you see evacuation failures, the first fix is more heap headroom, because the cause is usually a live set that has grown since the heap was sized. Only if the log shows concurrent cycles starting too late with plenty of headroom should you lower IHOP or raise -XX:G1ReservePercent, which keeps a percentage of the heap, 10 by default, free as a buffer. Change one flag at a time and compare logs from equivalent load.
Humongous objects
G1 divides the heap into equal regions, sized automatically from 1 MB to 32 MB in powers of two depending on heap size. An object at least half a region in size is humongous: it is allocated directly in old-generation regions of its own, contiguous if it spans several. With a 4 GiB heap the region size is typically 2 MB, so any array of 1 MB or more is humongous, such as a byte array buffering a large HTTP body or an ArrayList backing array that has grown large.
Many short-lived humongous allocations fragment the heap and can trigger concurrent cycles or full collections even when the live set is small. The log shows humongous regions in the heap summary. The fixes, in order: stream large payloads instead of buffering them, reuse or pool large buffers, chunk large arrays, and only then raise -XX:G1HeapRegionSize so those objects fall below half a region.
Choosing the collector
Heap tuning assumes a collector. G1 is the default on server-class machines and the right starting point for most services. The Parallel collector gives the highest throughput for batch jobs that tolerate long pauses. ZGC keeps pauses to around a millisecond almost independently of heap size, at a cost in throughput and memory headroom; it has been generational by default since JDK 23 and generational only since JDK 24. With ZGC, tuning is mostly setting -Xmx with enough headroom for concurrent collection to keep up, and optionally -XX:SoftMaxHeapSize. ZGC in depth and the garbage collection overview compare them.
OutOfMemoryError and other failure modes
OutOfMemoryError: Java heap space means the live set plus the garbage the collector could not reclaim in time exceeds -Xmx. It is either a heap that is too small for real load or a leak: a collection, cache or listener list that grows without bound. The heap dump answers which. Open it in a heap analyser, look at the dominator tree and follow the largest retained sizes back to the GC root that holds them. Two dumps taken an hour apart that show the same structure growing confirm a leak.
Do not try to keep running after a heap OOM. Threads may have died halfway through updating shared state, so the process is untrustworthy. -XX:+ExitOnOutOfMemoryError makes the JVM exit so the orchestrator restarts it, and the dump is still written first if you also set HeapDumpOnOutOfMemoryError.
| Symptom in the log or metrics | Usual cause | First action |
|---|---|---|
| Old occupancy after mixed GCs rising day over day | Leak or unbounded cache | Two heap dumps; compare dominators |
| Evacuation failure, then a full GC | Heap too small for the current live set | Remeasure live set; add headroom |
| Young pauses far over the goal | Large survivor copying, fixed young size, or CPU throttling | Remove -Xmn; check container CPU limits |
| Frequent concurrent cycles with low live set | Humongous allocations | Find large arrays with an allocation profiler |
| OutOfMemoryError: Metaspace | Class loader leak, not the heap | See metaspace guidance; -Xmx will not help |
| Container killed with no Java OOM | Process memory above the limit, not the heap | Native memory tracking; lower heap share |
A worked tuning session
A payments API on JDK 21 with G1 runs with -Xmx2g and reports p99 latency spikes at peak. The log shows young pauses every 1.5 seconds averaging 40 ms, mixed collections leaving 1.3 GiB, two evacuation failures per hour and a full collection lasting 1.8 seconds after each. The live set of 1.3 GiB is 65 percent of a 2 GiB heap, far below the three-to-four-times heuristic. The team sets -Xms5g -Xmx5g, nearly four times the live set, keeps the default pause goal and adds AlwaysPreTouch. Under the same load test, young pauses drop to every 4 seconds, evacuation failures and full collections disappear, and p99 latency falls back within budget. The container limit is raised to cover the heap plus the native memory measured with native memory tracking.
What to do next
- Enable rotated unified GC logging and heap dumps on OOM in every environment, with disk space for one dump.
- Capture a log at peak load and extract the live set after mixed collections, the allocation rate and p50 and p99 pauses.
- Size -Xmx at three to four times the live set, set -Xms to the same value and confirm container limits cover heap plus native memory.
- Remove -Xmn, NewRatio and other inherited flags from old guides, then set MaxGCPauseMillis from your latency budget.
- Check the log for evacuation failures, full collections and humongous regions, and fix each at its cause.
- Change one setting at a time, rerun the same load test and keep the logs from both runs for comparison.
- Add -XX:+ExitOnOutOfMemoryError, and alert on old-generation occupancy after mixed collections rising over days.