ForkJoinPool is the thread pool behind Java's parallel streams, the default executor for CompletableFuture's async methods, and, in a separate instance, the scheduler that runs virtual threads. Most developers use it without seeing it, until a parallel stream stops making progress, a service's async callbacks stall, or a CPU-bound job runs slower on sixteen cores than on four. Fixing those problems needs a picture of how the pool is built, not just its API.
This article opens the pool up. The ForkJoinPool basics article covers RecursiveTask, the common pool and ManagedBlocker, and the work-stealing scheduler article covers the deque algorithm in general. Here the focus is the pool's own structure: its queues, how idle workers are woken, what join() does before it blocks, when the pool adds threads, and how to read its state when something goes wrong. Internal names are from the OpenJDK implementation, which changes between releases, so treat them as a map rather than a contract.
Why a different pool at all
A classic ThreadPoolExecutor has one shared queue. Every submit and every take touches the same lock or the same atomic head, which is fine for a few thousand coarse tasks per second and a bottleneck for millions of tiny ones. Divide-and-conquer code produces exactly that: a task splits into two, each splits again, and a single job creates tens of thousands of subtasks.
ForkJoinPool gives every worker its own double-ended queue. A worker pushes and pops subtasks at one end without contention, because it is the only thread using that end. Idle workers steal from the other end of someone else's queue, which happens rarely because the oldest task in a divide-and-conquer queue is usually the largest, so one steal transfers a lot of work. The result is that most operations are uncontended and cheap, and load balances itself without a central dispatcher.
The queue array: submission slots and worker slots
The pool holds an array of work queues whose length is a power of two. In OpenJDK, queues fall into two kinds. Worker queues belong to one pool thread, which pushes and pops at the top; other threads only steal at the base, using compare-and-set. Submission queues receive tasks from threads outside the pool, such as your request thread calling pool.submit() or a parallel stream started from main. Those have no single owner, so pushes are guarded by a lightweight lock, and an external thread picks a submission queue by hashing a per-thread probe value, which spreads unrelated submitters across different queues.
In the implementation the two kinds have historically been interleaved, with submission queues on even indices and worker queues on odd ones, so a single scan over the array visits both. Workers look for work in a fixed order: their own queue first, then a scan of the array starting at a pseudo-random index, taking from whatever queue has tasks at its base. The random start keeps idle workers from all hitting the same victim.
The ctl word: counting and waking workers
A pool must decide when to start or wake a thread. Doing that with locks would reintroduce the contention the queues avoid, so the pool packs its control state into one 64-bit field, ctl, updated by compare-and-set. It holds counts of active and total workers relative to the target parallelism, and the top of a stack of idle workers. Idle workers form a linked stack through fields in their queues, with a version stamp to avoid the ABA problem; the stack top's index lives in ctl.
When a task is pushed onto a queue that may have been empty, the pushing thread calls an internal signalWork: if a worker is idle, pop it from the stack and unpark it; otherwise, if fewer workers exist than the parallelism target, create one. A worker that scans and finds nothing pushes itself onto the idle stack and parks, after a final rescan to avoid missing a task pushed in the meantime.
What join() really does
In a normal pool, a task that waits for another task blocks a thread. If all threads wait, the pool deadlocks. ForkJoinPool avoids this by making join() work instead of wait, in stages.
- If the joined task is already done, return its result.
- If it is still at the top of the caller's own queue, which is the common case for the last task it forked, pop it and run it directly. This is why the idiom forks one half and computes the other: the join usually finds its own task unstolen.
- If it was stolen, help: find the worker that stole it and run tasks from that worker's queue, which are subtasks of the stolen task, so helping makes progress on the very result being awaited.
- Only when helping runs out does the caller block. Before blocking, the pool may compensate: wake an idle worker or create a spare one, so that parallelism does not drop while this thread waits.
Compensation is bounded. The common pool's spares are limited by the maximumSpares property, and custom pools built with the extended constructor are limited by maximumPoolSize. When the limit is hit, a saturate predicate you supplied decides whether to continue without compensation; with no predicate, the pool throws RejectedExecutionException. Blocking I/O inside tasks therefore has two bad outcomes: thread counts grow up to the limit, or the pool's parallelism quietly drops to the number of threads not stuck in I/O. ForkJoinPool.managedBlock tells the pool in advance that a block is coming so compensation happens deliberately.
Worked example: a parallel histogram
Suppose you need a histogram of 200 million integers into 64 buckets. A sequential loop takes about a second on a typical machine. The task below splits the array until pieces are small enough, builds a private histogram per leaf, and merges results on the way back up. There is no shared mutable state, so no locks.
import java.util.concurrent.ForkJoinPool;
import java.util.concurrent.RecursiveTask;
/** Counts values per bucket in a large int array. Each leaf builds a private histogram. */
final class Histogram extends RecursiveTask<long[]> {
static final int THRESHOLD = 1 << 14; // ~16K elements: measure, then tune
final int[] data; final int lo, hi, buckets;
Histogram(int[] data, int lo, int hi, int buckets) {
this.data = data; this.lo = lo; this.hi = hi; this.buckets = buckets;
}
@Override protected long[] compute() {
if (hi - lo <= THRESHOLD) {
long[] h = new long[buckets];
for (int i = lo; i < hi; i++) h[Math.floorMod(data[i], buckets)]++;
return h;
}
int mid = (lo + hi) >>> 1;
Histogram left = new Histogram(data, lo, mid, buckets);
Histogram right = new Histogram(data, mid, hi, buckets);
left.fork(); // push onto this worker's own deque
long[] r = right.compute(); // keep working instead of waiting
long[] l = left.join(); // often pops 'left' back and runs it here
for (int b = 0; b < buckets; b++) r[b] += l[b];
return r;
}
}
// A dedicated pool keeps this CPU-bound job out of the shared common pool.
ForkJoinPool pool = new ForkJoinPool(Runtime.getRuntime().availableProcessors());
try {
long[] counts = pool.invoke(new Histogram(values, 0, values.length, 64));
} finally {
pool.shutdown();
}Choosing the threshold is the main tuning decision. Too large and there are fewer leaves than workers, so some cores idle at the end; too small and fork, join and stealing overhead dominate. A useful rule is to aim for leaves that take somewhere between tens of microseconds and a millisecond, then measure with a benchmark harness such as JMH. With 16,384 elements per leaf, 200 million elements give about 12,000 leaves, plenty for even a large machine.
When subtasks do not need to return values to a waiting parent, CountedCompleter removes joins altogether. Each task keeps a pending count; when a leaf finishes, it decrements its parent's count, and the last child to finish completes the parent. No thread ever waits, which is why the JDK's own parallel stream operations are built on it.
import java.util.concurrent.CountedCompleter;
import java.util.function.IntConsumer;
/** Applies an action to every index; nobody blocks in join(). */
final class ForEach extends CountedCompleter<Void> {
final int lo, hi; final IntConsumer action;
ForEach(CountedCompleter<?> parent, int lo, int hi, IntConsumer action) {
super(parent); this.lo = lo; this.hi = hi; this.action = action;
}
@Override public void compute() {
int l = lo, h = hi;
while (h - l > 4096) { // split off right halves as siblings
int mid = (l + h) >>> 1;
addToPendingCount(1);
new ForEach(this, mid, h, action).fork();
h = mid;
}
for (int i = l; i < h; i++) action.accept(i);
tryComplete(); // last finisher completes the parent chain
}
}
// new ForEach(null, 0, n, i -> out[i] = f(in[i])).invoke();
asyncMode and event-style workloads
By default, workers take their own tasks in last-in, first-out order, which suits recursion: the most recently forked task is the smallest and its data is still in cache. Tasks that are submitted and never joined suit FIFO better, because LIFO can starve old tasks. The asyncMode constructor argument switches local processing to FIFO. The virtual thread scheduler uses a ForkJoinPool in this mode, because each virtual thread continuation is an independent unit of work rather than a node in a recursion tree. The virtual threads article explains that model.
Configuration: common pool and custom pools
The common pool is created on first use with parallelism of available processors minus one, on the theory that the calling thread also helps. In container-aware JDKs, available processors reflects CPU limits from the container, but check what your JVM reports with Runtime.getRuntime().availableProcessors(), especially with fractional CPU quotas. A container reporting two or fewer processors gives the common pool parallelism of one, and CompletableFuture's async methods then create a new thread per task, which is a surprise in small containers.
# Common pool, set before the pool is first used (JVM flags):
-Djava.util.concurrent.ForkJoinPool.common.parallelism=4
-Djava.util.concurrent.ForkJoinPool.common.threadFactory=com.acme.NamedFactory
-Djava.util.concurrent.ForkJoinPool.common.exceptionHandler=com.acme.LogHandler
-Djava.util.concurrent.ForkJoinPool.common.maximumSpares=256
# The virtual thread scheduler is a separate ForkJoinPool with its own settings:
-Djdk.virtualThreadScheduler.parallelism=8For custom pools, the four-argument constructor takes parallelism, a thread factory, an uncaught exception handler and asyncMode. Since Java 9 an extended constructor also sets core pool size, maximum pool size, a minimum number of runnable threads, a saturate predicate and a keep-alive time for idle threads. Since Java 19, setParallelism changes the target at runtime.
Reading a pool's state
The pool's toString() packs most of what you need into one line, and the getters expose the same values for metrics. Thread dumps name common pool threads ForkJoinPool.commonPool-worker-N and custom pool threads ForkJoinPool-K-worker-N unless you supply a factory.
ForkJoinPool common = ForkJoinPool.commonPool();
System.out.println(common);
// java.util.concurrent.ForkJoinPool@1b6d3586[Running, parallelism = 7, size = 7,
// active = 7, running = 0, steals = 412, tasks = 0, submissions = 38]
//
// active 7, running 0, submissions growing: every worker is blocked inside a task.
System.out.printf("steals=%d queued=%d submissions=%d running=%d%n",
common.getStealCount(), common.getQueuedTaskCount(),
common.getQueuedSubmissionCount(), common.getRunningThreadCount());| Reading | Meaning | Next step |
|---|---|---|
| active = parallelism, running = 0, submissions rising | Every worker is blocked inside a task | Thread dump: look for I/O, locks or get() calls in worker stacks |
| size far above parallelism | Compensation has created spares for blocked joins | Remove blocking from tasks or use managedBlock deliberately |
| steals near zero on a parallel job | Work never split, or the root task is too coarse | Lower the threshold; check that fork is called |
| steals huge, speedup poor | Tasks too small; overhead dominates | Raise the threshold; measure leaf duration |
Failure modes and trade-offs
- Blocking in the common pool. A JDBC call inside
parallelStream()orsupplyAsyncwithout an executor ties up a shared worker. With parallelism seven, seven slow calls stall every parallel stream and async callback in the process. Give I/O its own executor or virtual threads. - Shared mutable state. Leaves that update a shared
HashMapor counter race or serialise on a lock. Give each leaf private state and merge on the way up, as the histogram does. - Exceptions.
join()andinvoke()rethrow a task's unchecked exception;get()wraps it inExecutionException. A forked task that is never joined swallows its exception silently. - Thread-locals and context. Tasks run on whichever worker takes them, so thread-local transaction, security or logging context from the submitting thread is absent. Pass context explicitly.
Choose ForkJoinPool for CPU-bound, recursively divisible work and for many small independent compute tasks. Choose a ThreadPoolExecutor for coarse tasks that need bounded queues, rejection policies and strict ordering; the ExecutorService article covers it. Choose virtual threads for blocking I/O at high concurrency. Mixing these roles in the common pool is the single most common cause of production trouble with it.
What to do next
- Search your codebase for parallelStream, supplyAsync and runAsync without an executor, and list which of them perform I/O or take locks.
- Move I/O-bound async work to a dedicated executor or virtual threads, and keep the common pool for CPU-bound work.
- Export getStealCount, getQueuedSubmissionCount, getActiveThreadCount, getRunningThreadCount and pool size as metrics for every pool you own.
- For each recursive task, measure leaf duration with JMH and tune the threshold so leaves take tens of microseconds to a millisecond.
- Give custom pools named thread factories and an uncaught exception handler, and shut them down when their job ends.
- Check availableProcessors inside your containers and set the common pool's parallelism explicitly if the reported value is wrong for your workload.