ForkJoinPool is the executor Java built for divide-and-conquer work. You split a problem into subtasks, solve them recursively and combine the results. What sets it apart from an ordinary thread pool is work-stealing. Every worker thread keeps its own queue of tasks, and a worker with nothing to do takes work from another worker's queue instead of waiting on one shared queue. That design keeps every core busy on recursive, CPU-bound work with very little coordination, and it is why parallel streams, Arrays.parallelSort and the async methods of CompletableFuture all run on a fork/join pool by default.

The same design is also the source of the most common production surprise in modern Java, which is a shared pool that silently stalls because something put blocking I/O on it. This article covers the programming model, then the internals that explain the performance: deques, stealing, helping joins and compensation threads. It then covers the common pool, how to block safely when you must, how exceptions and cancellation behave, and a worked diagnosis of common-pool starvation in a web service. It ends with a checklist. For the stream-level view, see Java parallel streams.

The fork/join model and the compute idiom

A fork/join computation is a task that checks its input size. If the input is small, the task solves it directly. If it is large, the task splits it, schedules one half with fork(), computes the other half itself and then calls join() to collect the forked result. Java gives you three task base classes. RecursiveTask<V> returns a value, RecursiveAction returns nothing, and CountedCompleter<T> triggers a completion action when a pending-count reaches zero.

import java.util.concurrent.*;

final class SumTask extends RecursiveTask<Long> {
    static final int THRESHOLD = 10_000;          // tune by measurement, see below
    final long[] a; final int lo, hi;
    SumTask(long[] a, int lo, int hi) { this.a = a; this.lo = lo; this.hi = hi; }

    @Override protected Long compute() {
        if (hi - lo <= THRESHOLD) {                // base case: plain loop
            long s = 0;
            for (int i = lo; i < hi; i++) s += a[i];
            return s;
        }
        int mid = (lo + hi) >>> 1;
        SumTask left  = new SumTask(a, lo, mid);
        SumTask right = new SumTask(a, mid, hi);
        left.fork();                               // push left onto this worker's deque
        long r = right.compute();                  // do right ourselves, no scheduling cost
        return left.join() + r;                    // left was probably not stolen: pop and run it
    }
}

long total = ForkJoinPool.commonPool().invoke(new SumTask(data, 0, data.length));

The ordering in that code matters. Fork one half and compute the other half in the current thread. Forking both halves and joining both leaves the current worker with nothing to do but wait. Join in the reverse order of your forks, newest first, because the most recently forked task is the one most likely to still sit on top of your own deque. ForkJoinTask.invokeAll(left, right) applies the right pattern for you and is a safe default when you have more than two subtasks.

Subtasks must be independent. They may read shared immutable input such as the array above, but they should not write shared mutable state. If you find yourself adding locks inside compute(), the decomposition is wrong.

Inside the pool: deques, stealing and helping joins

Each worker thread owns a double-ended queue. When the worker forks, it pushes onto the top of its own deque, and when it needs work it pops from the top. That makes the owner's queue behave like a stack. The task it runs next is the one it just created, whose data is still in cache, and the owner touches only its own end without locks in the common case.

A worker whose deque is empty becomes a thief. It scans other queues in a pseudo-random order and takes from the bottom, which holds the oldest task. In a recursive split the oldest task is the largest unsplit chunk, so one steal hands the thief a big piece of work that it will subdivide locally. Owner and thief also work opposite ends, so the two collide only when a deque is nearly empty, and that case is resolved with a compare-and-swap. Tasks submitted from threads that are not pool workers, such as a request thread calling invoke, land in separate submission queues that workers also scan.

Joins are where fork/join differs most from a future. A worker that calls join() on an unfinished task does not simply park. It first checks whether the task is still in its own deque and runs it directly. If the task was stolen, the worker tries to help by running tasks from the thief's queue, since that is the work standing between it and the result. Only when helping is impossible does it block, and then the pool may activate or create a compensation thread so that parallelism stays near its target. That is why a well-structured fork/join computation keeps every core busy despite thousands of joins.

Each worker owns a deque: it works LIFO at the top, thieves take FIFO from the bottomExternal callerpool.invoke(task)Submission queuesshared, for non-worker callerssubmitWorker 1deque: t9 t8 t5Worker 2deque: t7 t6Worker 3deque: emptyscanpush/pop topfork() pushes, compute()runs newest task: cache warmsteal oldestIdle worker stealsbottom of a victim dequejoin() helpsruns other tasks while waitingBlocked workerManagedBlocker lets the pool add a sparecommonPool()parallelism = cores - 1, shared JVM-wideNo central queue: load balances itself through steals, so contention stays low while tasks stay CPU-bound.
Workers push and pop their own deques LIFO, idle workers steal the oldest task from the other end, and joins help instead of waiting.

Choosing a split threshold

Splitting has a price. Each subtask is an object allocation, a deque push and possibly a steal. Too coarse a threshold leaves cores idle near the end of the computation, and too fine a threshold spends more time scheduling than computing. The ForkJoinTask Javadoc gives a rule of thumb: a task should perform more than about 100 and fewer than about 10,000 basic computational steps. The practical approach is to aim for several times more leaf tasks than worker threads, so stealing can even out uneven chunks, and then measure.

Two adaptive tricks are worth knowing. One is to compute a threshold from the input, for example Math.max(1_000, n / (pool.getParallelism() * 8)), which gives roughly eight leaves per worker. The other is getSurplusQueuedTaskCount(), which reports how many more tasks the current worker holds than other workers are likely to steal. Splitting only while that number is small, typically under 3, stops a worker from flooding its own deque when nobody is idle. Measure with JMH rather than a loop around System.nanoTime(), because JIT warm-up and dead-code elimination distort naive timing of exactly this kind of tight loop.

The common pool and who shares it

Since Java 8 every JVM has a static ForkJoinPool.commonPool(). Its target parallelism defaults to the number of available processors minus one, with a minimum of one, because the thread that submits work usually joins in. You can change it with the system property java.util.concurrent.ForkJoinPool.common.parallelism, and recent JDKs also offer setParallelism(int) on the pool. The common pool's threads are daemon threads, and calling shutdown() on it has no effect.

Several JDK APIs use the common pool unless you say otherwise. Parallel streams do. Arrays.parallelSort does. The CompletableFuture methods ending in Async do when you pass no executor, unless common-pool parallelism is 1 or less, in which case they start a new thread per task. Your own invoke calls do too if you use the common pool directly. Everything in the process shares one small set of threads sized for CPU work. One library doing a blocking call inside supplyAsync therefore takes a core away from every parallel stream in the application.

Virtual threads are a separate case. Their default scheduler is also a ForkJoinPool, running in FIFO mode with parallelism equal to the processor count, but it is a dedicated instance and not the common pool. Blocking a virtual thread unmounts it from its carrier, so the carrier pool does not starve the way the common pool does. Pinning cases are the exception; see Java virtual threads.

Blocking safely with ManagedBlocker

Fork/join assumes tasks never block. If a task must wait on a lock, a queue or I/O, wrap the wait in a ForkJoinPool.ManagedBlocker and call ForkJoinPool.managedBlock. The pool then knows a worker is about to park and may start a spare thread to keep parallelism up. The pattern below is adapted from the Javadoc:

final class QueueTaker<E> implements ForkJoinPool.ManagedBlocker {
    final BlockingQueue<E> queue;
    volatile E item;
    QueueTaker(BlockingQueue<E> q) { this.queue = q; }

    public boolean block() throws InterruptedException {
        if (item == null) item = queue.take();      // may park; pool can compensate
        return true;
    }
    public boolean isReleasable() {                  // checked first: avoid blocking at all
        return item != null || (item = queue.poll()) != null;
    }
}

QueueTaker<Job> taker = new QueueTaker<>(jobs);
ForkJoinPool.managedBlock(taker);
Job next = taker.item;

Compensation has limits. The common pool caps spare threads, configurable through java.util.concurrent.ForkJoinPool.common.maximumSpares with a default of 256. Spares also cost memory and context switches. ManagedBlocker is a safety valve for occasional, short waits, not a way to run an I/O workload on a CPU pool. For I/O fan-out, use virtual threads or a dedicated ExecutorService sized for the downstream dependency.

Custom pools, async mode, exceptions and monitoring

Building your own pool gives you control over the rest of the behaviour. The full constructor is new ForkJoinPool(parallelism, threadFactory, uncaughtExceptionHandler, asyncMode). Setting asyncMode to true makes each worker process its local tasks FIFO instead of LIFO. That suits event-style tasks that are forked but never joined, such as message handlers, and it is the mode the virtual-thread scheduler uses. For recursive decomposition, leave it false.

Exceptions behave differently by entry point. join() and invoke() rethrow a task's unchecked exception directly. get() wraps it in ExecutionException, like any future. When the exception crosses threads, the pool may rethrow a copy constructed in the joining thread with the original as its cause, so look at the cause chain for the real stack trace. cancel(true) ignores the interrupt flag, because fork/join tasks are not designed to be interrupted. Cancellation only prevents a task that has not started, so long-running leaves should check a shared flag or isCancelled() themselves.

For monitoring, the pool exposes getActiveThreadCount(), getRunningThreadCount(), getQueuedTaskCount(), getQueuedSubmissionCount() and getStealCount().

Worked example: diagnosing common-pool starvation

A checkout service had p99 latency jump from 120 ms to several seconds during a sale. Nothing obvious had changed. CPU sat at 35 percent. The trigger was a recent feature that enriched each cart line with a stock check. A developer had written cart.lines().parallelStream().map(stockClient::lookup), and lookup was a blocking HTTP call taking about 80 ms.

On a 16-core host the common pool has 15 workers. With 40 concurrent checkouts, each streaming ten lines, every worker was soon parked in a socket read. The service also built its pricing pipeline from CompletableFuture.supplyAsync(...) with no executor, and those tasks queued behind the stock checks even though they were pure CPU work. A thread dump made it plain: fifteen ForkJoinPool.commonPool-worker-N threads, all in SocketInputStream.read below ReferencePipeline frames, and a pool toString() showing hundreds of queued submissions with active threads equal to parallelism.

// Before: blocking I/O on the shared CPU pool
List<Stock> before = cart.lines().parallelStream().map(stockClient::lookup).toList();

// After: I/O on virtual threads, bounded by a semaphore that matches the stock service's capacity
try (var io = Executors.newVirtualThreadPerTaskExecutor()) {
    List<Future<Stock>> fs = cart.lines().stream()
        .map(l -> io.submit(() -> { permits.acquire();
                                    try { return stockClient.lookup(l); }
                                    finally { permits.release(); } }))
        .toList();
    List<Stock> s = new ArrayList<>();
    for (Future<Stock> f : fs) s.add(f.get(500, TimeUnit.MILLISECONDS));
}

// And: give CPU-bound async stages an explicit pool instead of the implicit common one
static final ForkJoinPool PRICING = new ForkJoinPool(Runtime.getRuntime().availableProcessors());
CompletableFuture.supplyAsync(() -> price(cart), PRICING);

After the change, the common pool went back to running only short CPU work, p99 returned to 130 ms, and the stock service saw a bounded concurrency of 64 instead of bursts it could not absorb. The team added two guards. One was a metric exporting the common pool's queued-submission count. The other was an ArchUnit rule that flags parallelStream() calls in classes that depend on the HTTP client package.

Failure modes and trade-offs

SymptomLikely causeFix
Parallel code slower than sequentialThreshold too small, tiny input, or a source that splits badly (linked list, iterator)Raise the threshold, keep small inputs sequential, use array-backed sources
Unrelated async work stallsBlocking calls on the common poolMove I/O to virtual threads or a dedicated executor; pass explicit executors
Thread count climbs into hundredsMany managed blocks or blocking joins triggering compensationRemove blocking from tasks; restructure with CountedCompleter
StackOverflowError in deep recursionJoins executing tasks inline on a deep split treeBalance splits by halving; avoid one-element-at-a-time decomposition
Wrong totalsShared mutable state written by subtasksReturn values and combine; never share accumulators
Results differ between runsNon-associative combine, such as floating-point sumsAccept tolerance, or use compensated summation and fixed splits

The trade-off is simple to state. Fork/join is the best general tool in the JDK for CPU-bound, splittable work on in-memory data, and a poor one for anything that waits. If the work is I/O-bound, use virtual threads with structured concurrency. If it is a stream of independent jobs, use a plain executor. If the per-element work is cheap and the data is small, stay sequential.

What to do next

  1. Search your code for parallelStream(), .parallel() and supplyAsync(/runAsync( calls without an executor, and confirm that each one does only CPU work.
  2. Add a metric or health endpoint that exports the common pool's parallelism, active count and queued-submission count, and alert when submissions keep growing.
  3. For each real fork/join task, check the idiom: fork one, compute one, join newest-first, or use invokeAll.
  4. Benchmark thresholds with JMH at production input sizes, aiming for several leaves per worker.
  5. Wrap any unavoidable wait inside a task in a ManagedBlocker, and plan to remove it.
  6. Give latency-critical CPU stages their own pool, so one noisy library cannot starve them.
  7. Read CompletableFuture in depth next to see how async pipelines pick their executors.
Key takeaway: ForkJoinPool keeps cores busy on recursive CPU work through per-worker deques, LIFO local execution, FIFO stealing and joins that help instead of waiting. Use it with the fork-one, compute-one idiom and a measured threshold. Treat the common pool as a shared, CPU-only resource, pass explicit executors to async APIs, and move anything that blocks to virtual threads or a dedicated executor.