Parallel GC, also called the throughput collector, is HotSpot's answer to one question: how do you spend as little total CPU time as possible on garbage collection? Its answer is to stop every application thread and use every available core to collect as fast as possible, then get out of the way. That makes it the cheapest collector per byte reclaimed, at the cost of pauses that grow with the heap.
It was the default server collector until JDK 9 and remains a strong choice for batch jobs, data pipelines and build tools. This article covers its heap layout, young and full collections, ergonomic goals, flags and logs (captured from JDK 23.0.2), then a worked batch-job example and trade-offs. For the shared vocabulary of roots, barriers and generations, see JVM GC architecture.
Heap layout and the card table
Parallel GC splits the heap into two contiguous generations, which the logs call PSYoungGen and ParOldGen. The young generation is an eden space where new objects are allocated, plus two equal survivor spaces, from and to. Application threads allocate from thread-local allocation buffers (TLABs) carved out of eden, so most allocations are a pointer bump with no locking.
Unlike G1, there are no regions here: each generation is a single reserved address range whose committed size can grow or shrink inside it. By default the old generation may be up to twice the size of the young generation (NewRatio=2), and each survivor space starts at one eighth of the young generation (InitialSurvivorRatio=8). The JDK 23 log below shows exactly that: 64,512 KB of eden and two 10,752 KB survivors. With adaptive sizing on, which it is by default, these are only starting points.
Old-to-young references are tracked with a card table, one byte for every 512 bytes of old generation. When the application stores a reference into an old object, a cheap write barrier marks that object's card dirty. At the next young collection the GC scans only dirty cards to find old objects that point into the young generation, instead of scanning the entire old generation.
The young collection: parallel scavenge
A young collection, sometimes called a scavenge, starts when eden is full. All application threads stop at a safepoint, and the GC worker threads split the work. Each worker takes a share of the roots: thread stacks, JNI handles, class metadata, and stripes of the card table's dirty cards. For every reachable young object it finds, the worker copies the object, either into the to-survivor space if the object is still young or into the old generation if its age has reached the tenuring threshold, and leaves a forwarding pointer behind so other references to it are updated to the new copy.
Copying is parallel through two mechanisms. Each worker copies into its own promotion-local allocation buffers (PLABs) in the survivor and old spaces, so workers rarely contend for space. Each worker also keeps a queue of objects whose fields still need scanning, and an idle worker steals work from another's queue, which keeps all cores busy even when one root leads to a huge object graph. When every queue is empty the phase ends and the survivor roles swap.
// Young collection, simplified pseudocode (all workers run this concurrently)
worker(id):
for root in claim_roots(id): # stacks, handles, dirty card stripes
root.ref = evacuate(root.ref)
while (obj = pop_local() or steal_from_others()):
for field in obj.reference_fields():
field.ref = evacuate(field.ref)
evacuate(o):
if o not in young: return o
if o.is_forwarded(): return o.forwardee()
dest = (o.age >= tenuring_threshold) ? old_plab : survivor_plab
if dest.full(): dest = refill_or_promote(o) # survivor overflow promotes early
copy = dest.allocate_and_copy(o); copy.age += 1
if not o.cas_forward(copy): return o.forwardee() # another worker won the race
push_local(copy); return copyThe cost of a young collection is proportional to the live objects in the young generation, not to its size, because dead objects are never touched. The danger is survivor overflow: if the survivor space cannot hold everything that survived, the excess is promoted to the old generation early, the old generation fills faster, and full collections become more frequent.
The full collection: parallel mark-compact
A full collection collects the whole heap and compacts the old generation, all in parallel and all with the application stopped. It runs when the old generation cannot absorb the next round of promotions, when metaspace needs to be collected, or when the application calls System.gc(). On JDK 23 the log shows its phases by name:
- Marking. Workers trace from the roots and mark every live object in both generations, recording how many live bytes fall into each fixed-size chunk of the heap.
- Summary. A quick pass over those per-chunk totals computes where each chunk's live data will go. It also chooses a dense prefix: a leading part of the old generation that is already almost entirely live, and so is not worth moving. Its few dead gaps are filled with filler objects instead.
- Forward. Each live object outside the dense prefix gets its new address.
- Adjust Pointers. Every reference in the heap and in the roots is rewritten to the new addresses.
- Compaction. Objects slide toward the start of the old generation, in parallel across chunks, preserving their order. Live young objects move into the old generation too, so young starts empty.
- Post Compact. Bookkeeping and resizing.
The cost of a full collection grows with both the live set, through marking and moving, and the heap size, through the per-chunk bookkeeping and pointer updates. By default UseMaximumCompactionOnSystemGC is true, so explicit System.gc() calls compact everything and ignore the dense prefix. The full-collection algorithm has been reworked upstream in recent years, so phase names differ between JDK versions; read the logs from the JDK you actually run.
Ergonomics: pause, throughput and footprint goals
Parallel GC can size the heap for you around goals rather than fixed sizes. There are three, and the collector applies them in a strict order. First comes the pause-time goal, -XX:MaxGCPauseMillis, which is not set by default for Parallel. If it is set, the collector shrinks generations when pauses run longer than the goal. Second comes the throughput goal, -XX:GCTimeRatio=N, which asks for at most 1/(1+N) of total time in GC. The default is 99, or 1 percent. When GC time exceeds the goal, the collector grows the generations. Third comes the footprint goal: only when both earlier goals are met does it shrink the heap toward the minimum.
The machinery behind this is the adaptive size policy (UseAdaptiveSizePolicy, on by default). After each collection it updates running averages of pause times, GC cost and promotion volumes, and adjusts eden, survivor and old sizes in steps; growth steps default to 20 percent (YoungGenerationSizeIncrement and TenuredGenerationSizeIncrement), and shrinking happens in smaller steps. It also adapts the tenuring threshold, starting at 7 and capped at 15, so objects are promoted sooner if the survivors keep overflowing.
In practice a 1 percent GC-time goal pushes the heap toward its maximum on any busy service, so -Xmx is effectively the heap size. Either let ergonomics work and tune the goals, or fix the sizes and turn the policy off.
How you get Parallel GC
Since JDK 9 the JVM never picks Parallel by itself. Through JDK 26, ergonomics chose G1 on machines it considered server class and Serial on small ones. JDK 27, released in September 2026, implements JEP 523, which makes G1 the default in all environments. Neither change selects Parallel, so you only ever get it by asking with -XX:+UseParallelGC. The old option of pairing a parallel young collector with a single-threaded old collector was deprecated in JDK 14 by JEP 366 and has since been removed, so the parallel compacting old collection is simply part of Parallel GC.
$ java -XX:+UseParallelGC -XX:+PrintFlagsFinal -version | grep -E 'GCTimeRatio|ParallelGCThreads|UseAdaptiveSizePolicy'
uint GCTimeRatio = 99 {product} {default}
uint ParallelGCThreads = 23 {product} {default}
bool UseAdaptiveSizePolicy = true {product} {default}The worker count is derived from the CPUs the JVM can see: one thread per CPU up to 8, then 8 plus five eighths of the rest. That is why the 32-CPU machine above gets 23 workers. In a container the JVM counts the CPU limit, not the host's cores.
The flags that matter
| Flag | Default (JDK 23) | Use it to |
|---|---|---|
-XX:+UseParallelGC | off | Select the collector; required |
-Xms / -Xmx | 1/64 and 1/4 of RAM | Fix the heap; set equal in containers to avoid resizing |
-XX:GCTimeRatio | 99 (1 percent) | Relax to 19 (5 percent) for a smaller heap, tighten for throughput |
-XX:MaxGCPauseMillis | unset | Ask for shorter pauses; trades throughput, not a guarantee |
-XX:ParallelGCThreads | derived from CPUs | Cap workers when several JVMs share a host |
-XX:-UseAdaptiveSizePolicy | on | Freeze sizes when you size generations by hand |
-Xmn | ergonomic | Fix young size, for example for a very high allocation rate |
-XX:+AlwaysPreTouch | off | Touch heap pages at startup so the first GCs do not page-fault |
Parallel also enforces a GC overhead limit by default: if almost all time goes to GC while very little heap is recovered, the JVM throws OutOfMemoryError: GC overhead limit exceeded rather than thrashing forever.
Reading the log
Enable logging with -Xlog:gc*:file=gc.log:time,uptime,level,tags:filecount=10,filesize=50m. These lines come from a small allocation test on JDK 23.0.2 with a 256 MB heap:
[0.013s][info][gc ] Using Parallel
[1.506s][info][gc,start] GC(1) Pause Young (Allocation Failure)
[1.510s][info][gc,heap ] GC(1) PSYoungGen: 67872K(75264K)->4624K(76288K) Eden: 64512K(64512K)->0K(65536K) From: 3360K(10752K)->4624K(10752K)
[1.510s][info][gc,heap ] GC(1) ParOldGen: 16K(172032K)->24K(172032K)
[1.510s][info][gc ] GC(1) Pause Young (Allocation Failure) 66M->4M(242M) 4.046ms
[1.960s][info][gc,phases] GC(44) Marking Phase 3.471ms
[1.960s][info][gc,phases] GC(44) Summary Phase 0.018ms
[1.962s][info][gc,phases] GC(44) Forward 1.498ms
[1.964s][info][gc,phases] GC(44) Adjust Pointers 1.834ms
[1.965s][info][gc,phases] GC(44) Compaction Phase 0.847ms
[1.965s][info][gc,heap ] GC(44) ParOldGen: 70823K(172032K)->2667K(172032K)
[1.965s][info][gc ] GC(44) Pause Full (System.gc()) 94M->2M(249M) 8.465msIn a young line, eden empties, survivors hold what lived and old grows by what was promoted; GC(1) promoted only 8 KB. The number to chart is old-generation occupancy after each full collection: that is your live set. A full collection with cause Ergonomics means the policy predicted the next promotion would not fit; frequent ones mean the old generation is too small for the promotion rate.
Worked example: a nightly ETL job
Here is a worked example. A nightly plain-Java ETL job reads 400 GB of records, transforms them and writes Parquet. It runs in a 16-CPU, 32 GB container and today uses G1 with defaults, finishing in 71 minutes. GC logs show about 9 percent of wall time in GC, plus concurrent-marking threads competing with the job for CPU.
First measure the live set. Old-generation occupancy after full collections peaks at about 6 GB, and the allocation rate is about 1.5 GB per second of short-lived record objects. The team switches to -XX:+UseParallelGC -Xms24g -Xmx24g -XX:NewRatio=1 -XX:+AlwaysPreTouch, leaving 8 GB for native memory. They keep adaptive sizing on. The default NewRatio of 2 would cap the young generation at a third of the heap, 8 GB, so NewRatio=1 lets it reach 12 GB, with eden around 9 GB once survivors are carved out. At 1.5 GB per second that means a young collection about every six seconds, ten per minute, each copying only the few hundred megabytes in flight. Pauses measure about 120 milliseconds, roughly 1.2 seconds of GC per minute, or 2 percent of wall time. Two full collections per run take about 2.5 seconds each.
The job now finishes in 63 minutes. They will revisit G1 if the job ever serves interactive traffic.
Failure modes
- Long full-GC pauses on large heaps. Pause length grows with the live set and the heap; several seconds is normal at tens of gigabytes. If any caller has a timeout, choose G1 or ZGC instead.
- Premature promotion. Survivor overflow pushes medium-lived objects into old, causing frequent full collections. Look for rising old occupancy after young GCs, and give survivors or eden more room.
- Undersized old generation. Back-to-back
Pause Full (Ergonomics)events with little reclaimed. Raise-Xmxor reduce the live set. - Oversubscribed host. Many JVMs collecting at once with full worker sets. Cap
ParallelGCThreads. - Explicit System.gc(). Libraries or RMI calls trigger full compactions. Find them by the cause in the log; consider
-XX:+DisableExplicitGConly after checking that nothing, such as direct-buffer cleanup, relies on them.
Trade-offs against Serial, G1 and ZGC
| Collector | Best at | Costs |
|---|---|---|
| Serial | One CPU, small heaps, lowest footprint | Single-threaded pauses, which grow fast with heap size |
| Parallel | Maximum throughput on many cores; batch and offline work | Full-heap stop-the-world pauses that grow with the live set |
| G1 | Balanced default; pause goals on medium and large heaps | Remembered sets, more barrier work, concurrent CPU use |
| ZGC | Sub-millisecond pauses at any heap size | Load barriers and higher CPU and memory overhead |
Choose Parallel when total work matters more than any single pause, the machine has several cores, and you can tolerate occasional pauses of hundreds of milliseconds to seconds. Continue with Serial GC, G1, reading GC logs and heap tuning.
What to do next
- List the JVMs that are throughput-bound rather than latency-bound: batch jobs, ETL, builds, offline scoring.
- For one of them, enable
-Xlog:gc*and record GC time as a share of wall time under the current collector. - Measure the live set from old-generation occupancy after full collections, and size
-Xmxto at least three times that. - Run the same job with
-XX:+UseParallelGC, equal-Xmsand-Xmxand adaptive sizing on; compare total runtime, GC share and peak RSS. - If several JVMs share a host, set
-XX:ParallelGCThreadsso their total is at or near the core count. - Pin the collector explicitly in every image, since JDK 27 changed the default for small environments, and write down why you chose it.