A thread pool is a fixed set of worker threads that take tasks from a queue. It exists because creating a platform thread costs a kernel thread, a stack reservation and scheduling overhead, so a server that starts one thread per request runs out of memory or drowns in context switches under load. Reusing threads caps that cost. It also creates a second, less obvious job: the pool becomes the place where your service decides how much work to accept, how much to queue and what to refuse.
This article is about that second job. The API surface, submit versus execute, the constructor parameters, rejection policies and shutdown, is covered in Java ExecutorService. Here we look at how a pool admits work, how to size it from measurements rather than folklore, which queue to choose, the deadlock pools create by themselves, and what to monitor so that saturation shows up on a dashboard before it shows up as timeouts. Examples use Java's ThreadPoolExecutor, but the reasoning applies to any pool.
The admission order, and the surprise inside it
When you call execute(task) on a ThreadPoolExecutor, it applies three rules in order. If fewer than corePoolSize threads are running, it starts a new thread with this task as its first job, even if other threads are idle. Otherwise it offers the task to the work queue. Only if the queue refuses the offer does it try to start a thread beyond core, up to maximumPoolSize. If that also fails, the task goes to the rejection handler.
Many engineers expect the opposite order: grow to the maximum, then queue. The real order has a consequence worth memorising. With an unbounded LinkedBlockingQueue, offer never fails, so the pool never grows past corePoolSize and maximumPoolSize is dead configuration. Executors.newFixedThreadPool(n) is built exactly this way, core equal to max and an unbounded queue, so under overload its queue grows without limit until the heap is exhausted or latency becomes meaningless.
The opposite extreme is Executors.newCachedThreadPool(): core zero, max Integer.MAX_VALUE, and a SynchronousQueue that holds nothing and accepts an offer only if a thread is already waiting to take it. Every burst creates threads, so a slow downstream can push it to thousands of threads. Neither factory has a bound on the thing that actually hurts, which is why production pools are usually constructed directly.
Choosing the queue
| Queue | Behaviour | Use when |
|---|---|---|
Unbounded LinkedBlockingQueue | Never rejects; max size ignored; latency grows without limit under overload | Tasks are internal, bounded in number by construction |
Bounded ArrayBlockingQueue or LinkedBlockingQueue(cap) | Fills, then grows threads to max, then rejects | Request-serving pools: the default choice |
SynchronousQueue | Direct hand-off; grows threads immediately; rejects at max | Short tasks where queueing adds no value and max is small |
PriorityBlockingQueue | Unbounded, ordered by priority; low priority can starve | Rarely; prefer separate pools per priority |
A bounded queue turns overload into a fast, visible signal, a rejection, instead of a slow invisible one, a queue that keeps growing. Size the queue from the latency you can tolerate, not from memory. If a pool of 20 threads completes 400 tasks per second and the caller's timeout budget leaves 250 ms for waiting, a queue longer than about 400 x 0.25 = 100 tasks only holds work that will time out before it runs. Work that will be abandoned anyway is better rejected at the door, where the caller can retry elsewhere or shed load. The general pattern of bounded hand-off between producers and consumers is covered in Java BlockingQueue.
For the rejection itself, CallerRunsPolicy is a useful form of back-pressure for internal pipelines: the submitting thread runs the task, so it cannot submit more until it finishes. It is a poor choice for a pool fed by a request thread that holds a socket, because it moves the slowness onto the very thread you wanted to protect. For request-serving pools, rejecting and returning a retryable error is usually more honest.
Sizing from first principles
Two formulas cover most sizing decisions. The first, popularised by Java Concurrency in Practice, estimates how many threads keep the CPUs busy:
threads = cores x target_cpu_utilisation x (1 + wait_time / compute_time)For purely CPU-bound work, wait time is zero and the answer is about the number of cores; more threads only add context switching. For work that spends most of its time blocked on I/O, the ratio dominates: a task that computes for 10 ms and waits 90 ms on a database needs about ten threads per core to keep that core busy.
The second is Little's law, which holds for any stable queueing system: the average number of items in the system equals the arrival rate times the average time each spends in it, L = lambda x W. For a pool, the number of tasks in flight equals throughput times task duration. It tells you the minimum concurrency you need, and it works backwards too: if you can only afford 20 concurrent tasks and each takes 100 ms, the pool cannot exceed 200 tasks per second however you tune it.
The two formulas answer different questions. The first asks how many threads saturate the CPU; Little's law asks how many concurrent tasks the traffic requires. Take the larger of the two as a starting point, then cap it by every downstream resource the tasks hold, which is the step most sizing exercises skip.
Worked example: an order-lookup service
A service runs on an 8-core machine. Each request does about 10 ms of CPU work (JSON parsing, business rules) and waits about 90 ms on one database query. Peak traffic is 400 requests per second. The database connection pool has 30 connections.
- CPU formula at 80 percent target utilisation: 8 x 0.8 x (1 + 90/10) = 64 threads.
- Little's law at peak: 400 per second x 0.1 s = 40 tasks in flight on average.
- CPU check: 400 x 10 ms = 4 core-seconds per second, half of the 8 cores. The CPU is not the limit.
- Downstream check: each in-flight task holds a connection for 90 ms, so 400 x 0.09 = 36 connections are needed on average. The pool has 30. The database connection pool, not the thread pool, is the real bottleneck.
A 64-thread pool here would let 64 tasks start, 30 of which hold connections while 34 sit blocked inside the connection pool, holding thread stacks and request memory while doing nothing. Raising threads does not raise throughput; it moves the queue somewhere you cannot see it. The right decision is either to raise the connection pool, if the database can take it, or to size the thread pool near the connection count, about 32 threads, with a bounded queue of around 100 and a rejection path. Then measure: if the observed wait time per task differs from 90 ms, recompute, because the formula is only as good as the ratio you feed it.
ThreadPoolExecutor pool = new ThreadPoolExecutor(
32, 32, // core = max: size is set by the DB, not by bursts
60, TimeUnit.SECONDS,
new ArrayBlockingQueue<>(100), // about 250 ms of queue at 400 tasks/s
new ThreadFactoryBuilder().setNameFormat("order-lookup-%d").build(), // Guava; or your own factory
new ThreadPoolExecutor.AbortPolicy()); // caller maps rejection to HTTP 503 + Retry-After
pool.allowCoreThreadTimeOut(true); // shrink when idle
Starvation deadlock: the pool deadlocks itself
A task running in a pool submits a subtask to the same pool and waits for its result. With enough concurrent parents, every thread is a parent waiting for a child, and every child sits in the queue waiting for a thread. Nothing holds a lock, so a deadlock detector looking for lock cycles finds nothing, but the pool has stopped. It happens at the worst time, under load, because only then are all threads busy with parents at once.
// Deadlocks once 'pool' has N threads all running handle() at the same time.
Result handle(Request r) throws Exception {
Future<Profile> profile = pool.submit(() -> loadProfile(r.userId()));
Future<List<Order>> orders = pool.submit(() -> loadOrders(r.userId()));
return new Result(profile.get(), orders.get()); // parent thread blocks waiting for children
}Fixes, in order of preference: do not block inside pool tasks, and compose with CompletableFuture instead so the parent returns its thread; put dependent stages in different pools so parents and children never compete for the same threads; or, for recursive divide-and-conquer, use a work-stealing ForkJoinPool, whose join runs queued subtasks instead of only waiting. The general patterns for finding hung threads are in deadlock detection; a thread dump of a starved pool shows every worker parked in FutureTask.get.
Exceptions that disappear
A task passed to execute that throws kills its worker thread, which the pool replaces; the exception goes to the thread's uncaught-exception handler and usually appears in the log. A task passed to submit behaves differently. It is wrapped in a FutureTask, which catches the exception and stores it for Future.get. If nobody calls get, the failure is never seen. Even the afterExecute(Runnable r, Throwable t) hook receives t == null for submitted tasks, because from the pool's point of view the FutureTask completed normally.
Two habits prevent silent loss: use execute for fire-and-forget work so failures reach the handler, and, if you override afterExecute for logging, unwrap the future when t is null and r is a completed Future, calling get() to retrieve the stored exception. A related leak: pool threads live for a long time, so a ThreadLocal set by one task is visible to the next task on that thread. Clear request-scoped thread-locals in a finally block or in afterExecute.
Metrics that show saturation early
ThreadPoolExecutor exposes getActiveCount(), getPoolSize(), getQueue().size(), getCompletedTaskCount() and getLargestPoolSize(). Export them, but know that queue depth alone is a lagging signal. The most useful metric is the time each task waits in the queue before it starts, and the pool does not record it. Wrap tasks to measure it:
final class TimedTask implements Runnable {
private final Runnable inner;
private final long enqueuedNanos = System.nanoTime();
TimedTask(Runnable inner) { this.inner = inner; }
@Override public void run() {
long waited = System.nanoTime() - enqueuedNanos;
QUEUE_WAIT.record(waited, TimeUnit.NANOSECONDS); // e.g. a Micrometer Timer
long start = System.nanoTime();
try { inner.run(); }
finally { RUN_TIME.record(System.nanoTime() - start, TimeUnit.NANOSECONDS); }
}
}
pool.execute(new TimedTask(() -> handle(request)));Alert on queue wait p99 relative to the caller's timeout, on the rejection rate, and on active threads sitting at the maximum for minutes. Name threads through a ThreadFactory so thread dumps and profilers attribute time to the right pool. A healthy pool shows queue wait near zero most of the time; a pool whose queue wait tracks traffic is under-provisioned or blocked downstream, and the run-time metric tells you which.
Bulkheads: one pool per failure domain
A single shared pool couples every dependency. If one downstream service slows from 50 ms to 5 s, its tasks hold threads a hundred times longer and soon occupy the whole pool, and requests that never touch that service fail too. Giving each dependency its own small bounded pool, a bulkhead, confines the damage: the slow dependency's pool saturates and rejects, the others keep running. The cost is some idle capacity and more configuration. Keep the number of pools small, one per external dependency or class of work, and size each from its own Little's-law numbers.
Where virtual threads change the picture
Since JDK 21, Executors.newVirtualThreadPerTaskExecutor() creates a cheap virtual thread per task. Blocking I/O parks the virtual thread and frees its carrier, so the wait/compute ratio no longer drives thread count, and pooling virtual threads defeats their purpose. What does not go away is the downstream limit from the worked example: a million virtual threads still share 30 database connections. With virtual threads you limit concurrency explicitly, typically with a Semaphore around the scarce resource, rather than implicitly through pool size.
Two cautions. CPU-bound work gains nothing, since there are still only as many carriers as cores. And on releases before JDK 24, a virtual thread that blocks inside a synchronized block pins its carrier; JEP 491 in JDK 24 removed that limitation. Native calls still pin. See virtual threads for the mechanics and thread pool basics for a shorter overview.
What to do next
- List every pool in your service, including those created by libraries, and record core, max, queue type and capacity, and rejection policy.
- Replace any
newFixedThreadPoolornewCachedThreadPoolon a request path with a directly constructed pool that has a bounded queue. - For each pool, measure compute time and wait time per task, then compute the CPU-formula size and the Little's-law size for peak traffic.
- Cap each pool by the downstream resource its tasks hold, such as connections or permits.
- Size queues from the caller's timeout budget, and map rejections to a retryable error.
- Search for tasks that call
Future.geton work submitted to the same pool, and restructure them. - Add queue-wait and run-time timers, name pool threads, and alert on queue-wait p99.
- Split shared pools into bulkheads for dependencies that have failed independently in the past.