ForkJoinPool is the executor Java built for divide-and-conquer work. You split a problem into subtasks, solve them recursively and combine the results. What sets it apart from an ordinary thread pool is work-stealing. Every worker thread keeps its own queue of tasks, and a worker with nothing to do takes work from another worker's queue instead of waiting on one shared queue. That design keeps every core busy on recursive, CPU-bound work with very little coordination, and it is why parallel streams, Arrays.parallelSort and the async methods of CompletableFuture all run on a fork/join pool by default.
The same design is also the source of the most common production surprise in modern Java, which is a shared pool that silently stalls because something put blocking I/O on it. This article covers the programming model, then the internals that explain the performance: deques, stealing, helping joins and compensation threads. It then covers the common pool, how to block safely when you must, how exceptions and cancellation behave, and a worked diagnosis of common-pool starvation in a web service. It ends with a checklist. For the stream-level view, see Java parallel streams.
The fork/join model and the compute idiom
A fork/join computation is a task that checks its input size. If the input is small, the task solves it directly. If it is large, the task splits it, schedules one half with fork(), computes the other half itself and then calls join() to collect the forked result. Java gives you three task base classes. RecursiveTask<V> returns a value, RecursiveAction returns nothing, and CountedCompleter<T> triggers a completion action when a pending-count reaches zero.
import java.util.concurrent.*;
final class SumTask extends RecursiveTask<Long> {
static final int THRESHOLD = 10_000; // tune by measurement, see below
final long[] a; final int lo, hi;
SumTask(long[] a, int lo, int hi) { this.a = a; this.lo = lo; this.hi = hi; }
@Override protected Long compute() {
if (hi - lo <= THRESHOLD) { // base case: plain loop
long s = 0;
for (int i = lo; i < hi; i++) s += a[i];
return s;
}
int mid = (lo + hi) >>> 1;
SumTask left = new SumTask(a, lo, mid);
SumTask right = new SumTask(a, mid, hi);
left.fork(); // push left onto this worker's deque
long r = right.compute(); // do right ourselves, no scheduling cost
return left.join() + r; // left was probably not stolen: pop and run it
}
}
long total = ForkJoinPool.commonPool().invoke(new SumTask(data, 0, data.length));The ordering in that code matters. Fork one half and compute the other half in the current thread. Forking both halves and joining both leaves the current worker with nothing to do but wait. Join in the reverse order of your forks, newest first, because the most recently forked task is the one most likely to still sit on top of your own deque. ForkJoinTask.invokeAll(left, right) applies the right pattern for you and is a safe default when you have more than two subtasks.
Subtasks must be independent. They may read shared immutable input such as the array above, but they should not write shared mutable state. If you find yourself adding locks inside compute(), the decomposition is wrong.
Inside the pool: deques, stealing and helping joins
Each worker thread owns a double-ended queue. When the worker forks, it pushes onto the top of its own deque, and when it needs work it pops from the top. That makes the owner's queue behave like a stack. The task it runs next is the one it just created, whose data is still in cache, and the owner touches only its own end without locks in the common case.
A worker whose deque is empty becomes a thief. It scans other queues in a pseudo-random order and takes from the bottom, which holds the oldest task. In a recursive split the oldest task is the largest unsplit chunk, so one steal hands the thief a big piece of work that it will subdivide locally. Owner and thief also work opposite ends, so the two collide only when a deque is nearly empty, and that case is resolved with a compare-and-swap. Tasks submitted from threads that are not pool workers, such as a request thread calling invoke, land in separate submission queues that workers also scan.
Joins are where fork/join differs most from a future. A worker that calls join() on an unfinished task does not simply park. It first checks whether the task is still in its own deque and runs it directly. If the task was stolen, the worker tries to help by running tasks from the thief's queue, since that is the work standing between it and the result. Only when helping is impossible does it block, and then the pool may activate or create a compensation thread so that parallelism stays near its target. That is why a well-structured fork/join computation keeps every core busy despite thousands of joins.
Choosing a split threshold
Splitting has a price. Each subtask is an object allocation, a deque push and possibly a steal. Too coarse a threshold leaves cores idle near the end of the computation, and too fine a threshold spends more time scheduling than computing. The ForkJoinTask Javadoc gives a rule of thumb: a task should perform more than about 100 and fewer than about 10,000 basic computational steps. The practical approach is to aim for several times more leaf tasks than worker threads, so stealing can even out uneven chunks, and then measure.
Two adaptive tricks are worth knowing. One is to compute a threshold from the input, for example Math.max(1_000, n / (pool.getParallelism() * 8)), which gives roughly eight leaves per worker. The other is getSurplusQueuedTaskCount(), which reports how many more tasks the current worker holds than other workers are likely to steal. Splitting only while that number is small, typically under 3, stops a worker from flooding its own deque when nobody is idle. Measure with JMH rather than a loop around System.nanoTime(), because JIT warm-up and dead-code elimination distort naive timing of exactly this kind of tight loop.
The common pool and who shares it
Since Java 8 every JVM has a static ForkJoinPool.commonPool(). Its target parallelism defaults to the number of available processors minus one, with a minimum of one, because the thread that submits work usually joins in. You can change it with the system property java.util.concurrent.ForkJoinPool.common.parallelism, and recent JDKs also offer setParallelism(int) on the pool. The common pool's threads are daemon threads, and calling shutdown() on it has no effect.
Several JDK APIs use the common pool unless you say otherwise. Parallel streams do. Arrays.parallelSort does. The CompletableFuture methods ending in Async do when you pass no executor, unless common-pool parallelism is 1 or less, in which case they start a new thread per task. Your own invoke calls do too if you use the common pool directly. Everything in the process shares one small set of threads sized for CPU work. One library doing a blocking call inside supplyAsync therefore takes a core away from every parallel stream in the application.
Virtual threads are a separate case. Their default scheduler is also a ForkJoinPool, running in FIFO mode with parallelism equal to the processor count, but it is a dedicated instance and not the common pool. Blocking a virtual thread unmounts it from its carrier, so the carrier pool does not starve the way the common pool does. Pinning cases are the exception; see Java virtual threads.
Blocking safely with ManagedBlocker
Fork/join assumes tasks never block. If a task must wait on a lock, a queue or I/O, wrap the wait in a ForkJoinPool.ManagedBlocker and call ForkJoinPool.managedBlock. The pool then knows a worker is about to park and may start a spare thread to keep parallelism up. The pattern below is adapted from the Javadoc:
final class QueueTaker<E> implements ForkJoinPool.ManagedBlocker {
final BlockingQueue<E> queue;
volatile E item;
QueueTaker(BlockingQueue<E> q) { this.queue = q; }
public boolean block() throws InterruptedException {
if (item == null) item = queue.take(); // may park; pool can compensate
return true;
}
public boolean isReleasable() { // checked first: avoid blocking at all
return item != null || (item = queue.poll()) != null;
}
}
QueueTaker<Job> taker = new QueueTaker<>(jobs);
ForkJoinPool.managedBlock(taker);
Job next = taker.item;Compensation has limits. The common pool caps spare threads, configurable through java.util.concurrent.ForkJoinPool.common.maximumSpares with a default of 256. Spares also cost memory and context switches. ManagedBlocker is a safety valve for occasional, short waits, not a way to run an I/O workload on a CPU pool. For I/O fan-out, use virtual threads or a dedicated ExecutorService sized for the downstream dependency.
Custom pools, async mode, exceptions and monitoring
Building your own pool gives you control over the rest of the behaviour. The full constructor is new ForkJoinPool(parallelism, threadFactory, uncaughtExceptionHandler, asyncMode). Setting asyncMode to true makes each worker process its local tasks FIFO instead of LIFO. That suits event-style tasks that are forked but never joined, such as message handlers, and it is the mode the virtual-thread scheduler uses. For recursive decomposition, leave it false.
Exceptions behave differently by entry point. join() and invoke() rethrow a task's unchecked exception directly. get() wraps it in ExecutionException, like any future. When the exception crosses threads, the pool may rethrow a copy constructed in the joining thread with the original as its cause, so look at the cause chain for the real stack trace. cancel(true) ignores the interrupt flag, because fork/join tasks are not designed to be interrupted. Cancellation only prevents a task that has not started, so long-running leaves should check a shared flag or isCancelled() themselves.
For monitoring, the pool exposes getActiveThreadCount(), getRunningThreadCount(), getQueuedTaskCount(), getQueuedSubmissionCount() and getStealCount().
Worked example: diagnosing common-pool starvation
A checkout service had p99 latency jump from 120 ms to several seconds during a sale. Nothing obvious had changed. CPU sat at 35 percent. The trigger was a recent feature that enriched each cart line with a stock check. A developer had written cart.lines().parallelStream().map(stockClient::lookup), and lookup was a blocking HTTP call taking about 80 ms.
On a 16-core host the common pool has 15 workers. With 40 concurrent checkouts, each streaming ten lines, every worker was soon parked in a socket read. The service also built its pricing pipeline from CompletableFuture.supplyAsync(...) with no executor, and those tasks queued behind the stock checks even though they were pure CPU work. A thread dump made it plain: fifteen ForkJoinPool.commonPool-worker-N threads, all in SocketInputStream.read below ReferencePipeline frames, and a pool toString() showing hundreds of queued submissions with active threads equal to parallelism.
// Before: blocking I/O on the shared CPU pool
List<Stock> before = cart.lines().parallelStream().map(stockClient::lookup).toList();
// After: I/O on virtual threads, bounded by a semaphore that matches the stock service's capacity
try (var io = Executors.newVirtualThreadPerTaskExecutor()) {
List<Future<Stock>> fs = cart.lines().stream()
.map(l -> io.submit(() -> { permits.acquire();
try { return stockClient.lookup(l); }
finally { permits.release(); } }))
.toList();
List<Stock> s = new ArrayList<>();
for (Future<Stock> f : fs) s.add(f.get(500, TimeUnit.MILLISECONDS));
}
// And: give CPU-bound async stages an explicit pool instead of the implicit common one
static final ForkJoinPool PRICING = new ForkJoinPool(Runtime.getRuntime().availableProcessors());
CompletableFuture.supplyAsync(() -> price(cart), PRICING);After the change, the common pool went back to running only short CPU work, p99 returned to 130 ms, and the stock service saw a bounded concurrency of 64 instead of bursts it could not absorb. The team added two guards. One was a metric exporting the common pool's queued-submission count. The other was an ArchUnit rule that flags parallelStream() calls in classes that depend on the HTTP client package.
Failure modes and trade-offs
| Symptom | Likely cause | Fix |
|---|---|---|
| Parallel code slower than sequential | Threshold too small, tiny input, or a source that splits badly (linked list, iterator) | Raise the threshold, keep small inputs sequential, use array-backed sources |
| Unrelated async work stalls | Blocking calls on the common pool | Move I/O to virtual threads or a dedicated executor; pass explicit executors |
| Thread count climbs into hundreds | Many managed blocks or blocking joins triggering compensation | Remove blocking from tasks; restructure with CountedCompleter |
| StackOverflowError in deep recursion | Joins executing tasks inline on a deep split tree | Balance splits by halving; avoid one-element-at-a-time decomposition |
| Wrong totals | Shared mutable state written by subtasks | Return values and combine; never share accumulators |
| Results differ between runs | Non-associative combine, such as floating-point sums | Accept tolerance, or use compensated summation and fixed splits |
The trade-off is simple to state. Fork/join is the best general tool in the JDK for CPU-bound, splittable work on in-memory data, and a poor one for anything that waits. If the work is I/O-bound, use virtual threads with structured concurrency. If it is a stream of independent jobs, use a plain executor. If the per-element work is cheap and the data is small, stay sequential.
What to do next
- Search your code for
parallelStream(),.parallel()andsupplyAsync(/runAsync(calls without an executor, and confirm that each one does only CPU work. - Add a metric or health endpoint that exports the common pool's parallelism, active count and queued-submission count, and alert when submissions keep growing.
- For each real fork/join task, check the idiom: fork one, compute one, join newest-first, or use
invokeAll. - Benchmark thresholds with JMH at production input sizes, aiming for several leaves per worker.
- Wrap any unavoidable wait inside a task in a
ManagedBlocker, and plan to remove it. - Give latency-critical CPU stages their own pool, so one noisy library cannot starve them.
- Read CompletableFuture in depth next to see how async pipelines pick their executors.