ExecutorService separates what work runs from which thread runs it. You submit tasks; the executor decides when and on which thread they execute, how many run at once and what happens when there is too much work. That decision point is where most Java server overload incidents are won or lost: an executor with the wrong queue turns a traffic spike into an OutOfMemoryError, and one with the wrong rejection policy turns it into silently dropped work.

This article explains the API and its lifecycle, the exact admission algorithm inside ThreadPoolExecutor, what the Executors factory methods really build, how to size a pool, and how to shut one down. It then builds a production-ready pool step by step and covers the failure modes and the virtual-thread alternative. Code targets JDK 21; methods newer than Java 8 are marked.

The API and its lifecycle

Three interfaces form a ladder. Executor has one method, execute(Runnable). ExecutorService adds results and a lifecycle: submit returns a Future; invokeAll runs a collection and waits for all; invokeAny returns the first successful result and cancels the rest; and shutdown, shutdownNow and awaitTermination stop it. ScheduledExecutorService adds delayed and periodic execution.

JDK 19 added two conveniences. ExecutorService now extends AutoCloseable, and its close() shuts the executor down and waits for running and queued tasks to finish, so try-with-resources gives you structured clean-up. Future gained state(), resultNow() and exceptionNow() for inspecting completed futures without checked exceptions.

The difference between execute and submit matters for errors. An exception thrown by an executed Runnable kills that worker thread, reaches the thread's uncaught-exception handler and is usually logged; the pool replaces the worker. An exception thrown by a submitted task is captured inside its Future and appears only when someone calls get(). If nobody does, the failure is invisible. This is the single most common way background work fails silently.

How ThreadPoolExecutor admits a task

Almost every executor you meet is a ThreadPoolExecutor configured with a core size, a maximum size, a keep-alive time, a work queue, a thread factory and a rejection handler. Its execute method follows three steps, documented in its Javadoc:

  1. If fewer than corePoolSize workers exist, start a new one to run the task, even if other workers are idle.
  2. Otherwise, try to put the task in the queue.
  3. If the queue refuses, start another worker if that stays within maximumPoolSize; otherwise hand the task to the rejection handler.

ThreadPoolExecutor.execute(task): the admission order that surprises peopleexecute(task)workers < corePoolSize?yesstart a new workereven if others are idlenoqueue.offer(task) succeeds?yestask waits in the queuepool stays at core sizeno (queue full)workers < maximumPoolSize?yesstart an extra workerit runs this task firstnoRejectedExecutionHandlerabort, caller-runs, discard...ConsequencesUnbounded queue: step 3 neverruns, so max is ignored.SynchronousQueue: offer failsunless a worker is waiting.Extra workers above core exit after keepAliveTime idle; core workers too if allowCoreThreadTimeOut(true).A shut-down pool rejects new tasks through the same handler.
The admission decision inside ThreadPoolExecutor. The queue is tried before the pool grows past core.

The surprise is step 2 before step 3: the pool grows beyond core only when the queue is full. The queue type therefore decides how the pool behaves:

QueueBehaviourTypical use
SynchronousQueuehands each task directly to an idle worker or grows the pool; nothing waitsshort tasks with a hard thread cap
unbounded LinkedBlockingQueuepool never exceeds core; backlog grows without limitalmost never what you want on a server
bounded ArrayBlockingQueuecore threads, then a finite backlog, then extra threads, then rejectionthe default for server work

The bounded queue is the only one of the three with predictable worst-case memory. The BlockingQueue article explains how to size it.

What the Executors factories really build

The Executors factory methods are convenient, and several hide exactly the dangerous settings:

FactoryWhat it buildsRisk
newFixedThreadPool(n)core = max = n, unbounded LinkedBlockingQueuebacklog grows until OOM
newSingleThreadExecutor()one thread, unbounded queuesame, plus one slow task blocks all
newCachedThreadPool()core 0, max Integer.MAX_VALUE, SynchronousQueue, 60 s keep-alivea burst creates thousands of threads
newScheduledThreadPool(n)ScheduledThreadPoolExecutor with an unbounded internal delay queuemaximum size has no effect
newWorkStealingPool()a ForkJoinPool sized to the processorsunsuited to blocking tasks
newVirtualThreadPerTaskExecutor() (JDK 21)a new virtual thread per task, no poolunlimited concurrency against downstreams

Alibaba's Java development manual, for example, forbids the factories for exactly these reasons and asks for explicit ThreadPoolExecutor construction. You do not need a rule that strict, but every pool that handles external load should have a visible queue bound and a deliberate rejection policy. For CPU-heavy divide-and-conquer work, ForkJoinPool is the better tool.

Sizing a pool

Brian Goetz's sizing rule from Java Concurrency in Practice is the right starting point: threads = cores × target utilisation × (1 + wait time / compute time). A CPU-bound task has almost no wait time, so about one thread per core. An I/O-bound task waits most of the time, so more threads keep the cores busy.

Worked example. A service calls a pricing API. Each task spends 8 ms computing and 72 ms waiting for the network. On an 8-core machine targeting 80% utilisation: 8 × 0.8 × (1 + 72/8) = 64 threads. Check it against the downstream. If the pricing API allows 50 concurrent requests per client, 64 threads will exceed it, so cap the pool at 50 and let the queue absorb the difference. Then size the queue with Little's law: 50 threads finishing an 80 ms task give about 625 tasks per second. If the caller's deadline allows 200 ms of queueing, the bound is about 625 × 0.2 = 125.

Measure the two inputs rather than guessing them. Time each task's total duration and its CPU time, for example with ThreadMXBean.getCurrentThreadCpuTime() sampled at the start and end of the task, or read the split from a Java Flight Recorder profile. Wait time is the difference. Re-measure when a dependency changes, because a slower downstream raises the ideal thread count and the queue's latency at the same time.

Give each kind of work its own pool. This bulkhead pattern stops a slow dependency from consuming every thread in the process. A pool for pricing calls and a separate pool for image resizing fail independently.

Worked example: a pool for a pricing service

Here is a pool built for the pricing workload above. Every choice is explicit and every failure is visible:

public final class PricingPool {
    private final AtomicLong rejected = new AtomicLong();

    private final ThreadPoolExecutor pool = new ThreadPoolExecutor(
            50, 50,                                     // core = max: fixed size, capped by downstream
            30, TimeUnit.SECONDS,
            new ArrayBlockingQueue<>(125),              // ~200 ms of backlog at 625 tasks/s
            Thread.ofPlatform().name("pricing-", 0)     // JDK 21 builder; numbered thread names
                  .uncaughtExceptionHandler((t, e) -> log.error("uncaught in {}", t.getName(), e))
                  .factory(),
            (task, executor) -> {                       // fail fast and count it
                rejected.incrementAndGet();
                throw new RejectedExecutionException("pricing pool saturated");
            }) {
        @Override protected void afterExecute(Runnable r, Throwable t) {
            super.afterExecute(r, t);
            if (t == null && r instanceof Future<?> f && f.isDone()) {
                if (f.state() == Future.State.FAILED) {             // JDK 19+
                    log.warn("pricing task failed", f.exceptionNow());
                }
            }
            if (t != null) log.error("pricing task threw", t);
        }
    };

    { pool.allowCoreThreadTimeOut(true); }             // shrink to zero when idle

    public CompletableFuture<Price> quote(Sku sku) {
        try {
            return CompletableFuture.supplyAsync(() -> client.price(sku), pool);
        } catch (RejectedExecutionException e) {
            return CompletableFuture.failedFuture(new OverloadedException(e)); // caller maps to 503
        }
    }

    public int queueDepth() { return pool.getQueue().size(); }
    public int active()     { return pool.getActiveCount(); }
    public long rejected()  { return rejected.get(); }
}

The afterExecute hook catches failures of tasks handed to submit; a task run through CompletableFuture records its failure in the returned future instead, so the caller must handle it there. Trace a spike. The first 50 concurrent quotes each get a thread. The next 125 wait in the queue, adding at most about 200 ms. Quote 176 is rejected at once, the caller returns 503, and the rejection counter rises. Memory stays flat, the downstream never sees more than 50 concurrent calls, and the overload shows up in metrics rather than in the garbage collector. Because the task is passed explicitly to supplyAsync, it does not fall back to the shared ForkJoinPool.commonPool(), which is what happens when you omit the executor; see CompletableFuture.

Saturation and rejection policies

When the queue and the pool are both full, the rejection handler runs. The built-in policies:

PolicyBehaviourUse when
AbortPolicy (default)throws RejectedExecutionExceptioncallers can shed load, for example with a 503
CallerRunsPolicyruns the task on the submitting threadthe submitter is a producer that may safely slow down
DiscardPolicydrops the new task silentlyrarely; losses are invisible
DiscardOldestPolicydrops the oldest queued task and retriesonly the newest data matters, such as sensor readings

CallerRunsPolicy is often recommended as free back-pressure, and for batch producers it is. There are two catches. If the submitter is a request or event-loop thread, a slow task now runs on that thread and stalls everything behind it. And once the executor is shut down, CallerRunsPolicy discards the task instead of running it, so work submitted during shutdown is lost silently.

Shutting down cleanly

Scheduled executors deserve a word here. scheduleAtFixedRate keeps a fixed cadence and, if a run overruns, starts the next one late rather than concurrently; scheduleWithFixedDelay waits a fixed gap after each run ends. Use the second for polling a slow dependency, so a stall does not turn into back-to-back runs.

An executor's threads are non-daemon by default, so a pool you never shut down keeps the JVM alive. The standard two-phase shutdown:

pool.shutdown();                                   // refuse new tasks, finish queued ones
try {
    if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
        List<Runnable> unstarted = pool.shutdownNow();   // interrupt running tasks
        persistOrLog(unstarted);                         // queued work you are abandoning
        if (!pool.awaitTermination(10, TimeUnit.SECONDS))
            log.error("pool did not terminate");
    }
} catch (InterruptedException e) {
    pool.shutdownNow();
    Thread.currentThread().interrupt();
}

shutdownNow() only interrupts. A task that never checks its interrupt flag, or that swallows InterruptedException, keeps running. Write tasks that respond to interruption. On JDK 19 or later, short-lived executors can use try-with-resources; close() waits for all tasks, so use it where waiting for everything is the intended behaviour.

Failure modes

Most executor incidents fall into a few patterns:

  • Thread-starvation deadlock. A task submits a subtask to the same bounded pool and blocks on its get(). When every thread does this, nothing can run the subtasks. Use separate pools, or compose asynchronously without blocking.
  • Silent failure. Submitted tasks whose futures are never read. Inspect results, or log in afterExecute as above.
  • Periodic tasks that stop. If a scheduleAtFixedRate task throws, later executions are suppressed with no further warning. Catch everything inside periodic tasks.
  • ThreadLocal leakage. Pooled threads live for a long time, so a ThreadLocal set by one task is visible to the next. Clear it in a finally block.
  • Pools created per request. A method that builds an executor, submits one task and never shuts it down leaks threads on every call. Create pools once, at start-up, and own their shutdown.
  • The hidden unbounded queue. A factory pool under sustained overload. Heap usage climbs, GC pauses lengthen, latency grows, then the process dies.

Virtual-thread executors

Since JDK 21, Executors.newVirtualThreadPerTaskExecutor() starts a new virtual thread for every task. Virtual threads are cheap enough that pooling them is pointless, so for I/O-bound work this often replaces a carefully sized platform pool. Two things change. Concurrency is no longer limited by thread count, so protect downstreams explicitly, typically with a Semaphore of 50 permits around the pricing call. And CPU-bound work gains nothing; keep a platform pool near the core count for it.

Virtual threads that blocked inside synchronized used to pin their carrier thread. JEP 491 removed most of that pinning in JDK 24, so on JDK 21 prefer ReentrantLock on hot blocking paths. For fan-out with cancellation, look at StructuredTaskScope, and check its preview status on your JDK.

What to do next

To put this into practice:

  1. Find every Executors.new... call in your services and note its real queue and thread bounds.
  2. Replace pools that handle external load with explicit ThreadPoolExecutors that have bounded queues and deliberate rejection policies.
  3. Size each pool from measured wait and compute times, then cap it at the downstream's concurrency limit.
  4. Give threads meaningful names and an uncaught-exception handler, and log failed futures.
  5. Export queue depth, active threads and rejections for every pool, and alert on sustained rejection.
  6. Implement two-phase shutdown and test that tasks respond to interruption.
  7. For I/O-bound work on JDK 21 or later, trial a virtual-thread-per-task executor with a semaphore guarding each downstream.

Key takeaway: An ExecutorService is a policy for overload as much as a thread pool. Know the core, then queue, then max admission order, bound every queue, choose the rejection policy on purpose, size from wait and compute times capped by the downstream, make failures visible, shut down in two phases, and use virtual threads plus semaphores for I/O-bound work.