Most Java developers meet the Executor framework through three calls: a factory method on Executors, submit and shutdown. Underneath is a small, carefully engineered set of classes: a lifecycle state machine packed into one integer, worker threads that double as locks, a future that is a lock-free state machine of its own, and a scheduler built on a binary heap. Knowing that architecture explains behaviour that otherwise looks mysterious: why a thrown exception vanishes, why a periodic task silently stops, why shutdown hangs, and why maximumPoolSize does nothing on a scheduled pool.

This article walks the framework from its interfaces down to the internals of ThreadPoolExecutor as written in OpenJDK, then through FutureTask, the scheduled pool, completion services and the lifecycle, and ends with an instrumented executor you can drop into a service. Queue choice and pool sizing are covered in thread pools, and everyday usage in the ExecutorService guide.

The layers

The framework separates what to run from how to run it, in four layers:

  • Executor: one method, execute(Runnable). It promises nothing about threads; an implementation may even run the task in the caller.
  • ExecutorService: adds lifecycle (shutdown, shutdownNow, awaitTermination and, since Java 19, close through AutoCloseable) and result-bearing submission (submit, invokeAll, invokeAny).
  • AbstractExecutorService: implements submit and the invoke methods once, on top of execute, by wrapping each task in a RunnableFuture produced by the protected hook newTaskFor. Concrete classes only supply execute and the lifecycle.
  • Implementations: ThreadPoolExecutor; ScheduledThreadPoolExecutor, which extends it and implements ScheduledExecutorService; ForkJoinPool, which also extends AbstractExecutorService but runs tasks by work stealing; and the thread-per-task executor behind Executors.newVirtualThreadPerTaskExecutor(), final since Java 21.
Executorexecute(Runnable)ExecutorServicelifecycle, submit, invokeAbstractExecutorServicenewTaskFor wraps tasksThreadPoolExecutorctl, workers, queueScheduledThreadPoolExecutorDelayedWorkQueue heapexecute(task)core worker, else queue, else extra worker, else rejectctl3 state bits, 29 countBlockingQueuewaiting tasksRejection handlerabort, caller-runsCAS countofferrefusedWorkers: non-reentrant AQS locksrunWorker: getTask, beforeExecute, run, afterExecutetake or pollFutureTaskNEW, then NORMAL, EXCEPTIONAL or CANCELLEDrun()
Left: the type layering. Right: inside ThreadPoolExecutor, execute updates ctl, the queue or the rejection handler; workers pull tasks and run them as FutureTasks.

newTaskFor is the cleanest place to wrap every task, for example to carry tracing context across threads.

ctl: the lifecycle in one int

ThreadPoolExecutor keeps its run state and worker count in a single AtomicInteger named ctl. The top 3 bits hold the run state and the low 29 bits the worker count, so a pool can count at most 2^29 - 1 (about 536 million) workers. Packing both into one word lets the pool change the count and check the state in a single compare-and-set, which closes the race where a worker is added just as the pool shuts down.

private static final int COUNT_BITS = Integer.SIZE - 3;        // 29
private static final int COUNT_MASK = (1 << COUNT_BITS) - 1;

private static final int RUNNING    = -1 << COUNT_BITS;  // accepts and runs tasks
private static final int SHUTDOWN   =  0 << COUNT_BITS;  // no new tasks, drains queue
private static final int STOP       =  1 << COUNT_BITS;  // no new tasks, drops queue, interrupts
private static final int TIDYING    =  2 << COUNT_BITS;  // no workers left, terminated() next
private static final int TERMINATED =  3 << COUNT_BITS;  // terminated() has returned

private static int runStateOf(int c)     { return c & ~COUNT_MASK; }
private static int workerCountOf(int c)  { return c & COUNT_MASK; }
private static int ctlOf(int rs, int wc) { return rs | wc; }

Because RUNNING is negative and the states increase in order, 'is the pool at least SHUTDOWN' is a plain integer comparison. Transitions only move forward: RUNNING to SHUTDOWN on shutdown(); RUNNING or SHUTDOWN to STOP on shutdownNow(); to TIDYING once the worker count is zero (and, from SHUTDOWN, the queue is empty); and to TERMINATED when the protected terminated() hook returns. awaitTermination waits on a condition signalled at that last step.

execute and addWorker

execute has three steps: below corePoolSize, start a new worker with this task as its first task; otherwise offer the task to the queue; if the queue refuses, try to start a worker up to maximumPoolSize, and if that fails, hand the task to the RejectedExecutionHandler. The thread pools article explains what that order means for queue choice. Architecturally, two details matter.

First, adding a worker is two-phase. addWorker loops on a CAS that increments the count in ctl, re-checking run state and limits each time. Only after the CAS wins does it build the Worker, take the pool-wide mainLock to add it to the workers set, and start its thread, rolling the count back if the start fails. Queueing a task never touches that lock.

Second, after a successful queue offer, execute re-reads ctl. If the pool shut down in the meantime, it removes the task and rejects it; if the worker count is zero, it starts a worker with no first task so the queued task is not stranded. That recheck is why a pool with corePoolSize 0 still runs its queued work.

Workers: threads that are also locks

Each worker is an instance of the private class Worker, which extends AbstractQueuedSynchronizer and implements Runnable. It holds its thread, an optional first task and a completed-task counter, and its run method calls the pool's runWorker loop, shown here simplified:

final void runWorker(Worker w) {
    Runnable task = w.firstTask;
    w.firstTask = null;
    w.unlock();                        // state -1 to 0: interrupts now allowed
    boolean completedAbruptly = true;
    try {
        while (task != null || (task = getTask()) != null) {
            w.lock();                  // marks this worker busy
            // (if the pool is stopping, ensure this thread is interrupted)
            try {
                beforeExecute(w.thread, task);
                try {
                    task.run();
                    afterExecute(task, null);
                } catch (Throwable ex) {
                    afterExecute(task, ex);
                    throw ex;
                }
            } finally {
                task = null;
                w.completedTasks++;
                w.unlock();
            }
        }
        completedAbruptly = false;
    } finally {
        processWorkerExit(w, completedAbruptly);
    }
}

The lock is how the pool tells idle workers from busy ones. To wake idle workers, interruptIdleWorkers calls tryLock on each: success means the worker is blocked in getTask between tasks and safe to interrupt. The lock is deliberately non-reentrant, so a task that calls a control method such as setCorePoolSize cannot take its own worker's lock and interrupt itself. The initial state of -1 blocks interrupts until the thread is running.

getTask decides whether a worker lives. With more than corePoolSize workers, or with allowCoreThreadTimeOut set, it uses a timed poll(keepAliveTime); otherwise it blocks in take(). Returning null retires the worker after decrementing the count. A task that throws escapes runWorker, marks the exit as abrupt, and processWorkerExit replaces the dead worker with a fresh one: the pool survives throwing tasks but pays for a new thread each time.

FutureTask: a small lock-free state machine

submit wraps the task in a FutureTask, and that wrapper is why exceptions from submitted tasks disappear. FutureTask.run catches every Throwable and stores it as the outcome, so runWorker sees a normal return and afterExecute receives null. The exception resurfaces only when someone calls get, wrapped in ExecutionException.

Internally FutureTask holds a volatile int state with seven values: NEW, COMPLETING, NORMAL, EXCEPTIONAL, CANCELLED, INTERRUPTING and INTERRUPTED. The legal paths are NEW to COMPLETING to NORMAL or EXCEPTIONAL; NEW to CANCELLED; and NEW to INTERRUPTING to INTERRUPTED. A CAS on the runner field ensures only one thread runs the task, and a CAS on state settles races between completion and cancellation. Threads blocked in get push themselves onto a lock-free stack of wait nodes and park; completion pops and unparks them all.

cancel(true) moves the state to INTERRUPTING and interrupts the runner. That only stops the work if the task responds to interruption; a task in a tight CPU loop keeps running and its result is discarded. Java 19 added Future.state(), resultNow() and exceptionNow(), which inspect a completed future without the checked-exception ceremony of get. Composition on top of futures is covered in futures and promises.

ScheduledThreadPoolExecutor

The scheduled pool reuses ThreadPoolExecutor's workers but replaces the queue with a DelayedWorkQueue: a binary heap of ScheduledFutureTask objects ordered by trigger time, with a sequence number breaking ties so tasks due at the same instant run in submission order. A worker taking from the queue waits only until the head is due; one waiting thread acts as leader and waits with a timeout while the others wait indefinitely, which avoids a herd of timed waits.

Three consequences surprise people. The queue is unbounded, so the pool never grows past corePoolSize and maximumPoolSize has no useful effect. A periodic task that throws is never rescheduled, and the only evidence is the exception held in its future. And a cancelled task stays in the heap until its trigger time unless setRemoveOnCancelPolicy(true) is set.

scheduleAtFixedRate computes the next run from the previous scheduled time, so a slow run is followed by late runs that try to catch up; scheduleWithFixedDelay measures from the end of the previous run. Neither runs one task concurrently with itself.

Completion services and invokeAny

ExecutorCompletionService pairs an executor with a queue of completed futures. Each task is wrapped so that, when done, its future is added to the queue, and take() or poll return results in completion order rather than submission order. invokeAny in AbstractExecutorService is built on it: submit the tasks, take the first successful result, cancel the rest. invokeAll waits for every future and returns them in submission order; its timed form cancels whatever has not finished at the deadline.

var ecs = new ExecutorCompletionService<Quote>(pool);
List<Future<Quote>> pending = new ArrayList<>();
for (Supplier s : suppliers) pending.add(ecs.submit(() -> s.quote(item)));

long deadline = System.nanoTime() + TimeUnit.MILLISECONDS.toNanos(200);
Quote best = null;
for (int i = 0; i < pending.size(); i++) {
    Future<Quote> f = ecs.poll(deadline - System.nanoTime(), TimeUnit.NANOSECONDS);
    if (f == null) break;                       // the rest are too slow
    try { best = Quote.cheaper(best, f.get()); }
    catch (ExecutionException e) { log.warn("supplier failed", e.getCause()); }
}
pending.forEach(f -> f.cancel(true));           // no-op for finished ones

The final line matters: futures you stop waiting for keep running unless you cancel them, and their threads stay busy.

Shutdown and close

shutdown() stops admission, lets queued tasks finish and interrupts only idle workers. shutdownNow() drains the queue, returns the tasks that never started and interrupts every started worker. Neither waits, so the idiom is to shut down, wait with a bound, then escalate:

pool.shutdown();
if (!pool.awaitTermination(30, TimeUnit.SECONDS)) {
    List<Runnable> neverRan = pool.shutdownNow();
    log.warn("{} queued tasks dropped", neverRan.size());
    if (!pool.awaitTermination(10, TimeUnit.SECONDS))
        log.error("tasks ignore interruption; pool did not terminate");
}

Since Java 19 every ExecutorService is AutoCloseable. The default close calls shutdown and waits for termination with no time limit; if the waiting thread is interrupted, it calls shutdownNow and continues waiting. That makes try-with-resources a natural scope for a batch of tasks, and a hazard for a long-lived pool, because close blocks for as long as the slowest task. Tasks that never check interruption defeat every one of these methods; the pool can only ask.

Worked example: an instrumented pool

The protected hooks make it straightforward to build a pool that reports run time and the exceptions submit would otherwise swallow, two things missing from most incident timelines:

public final class InstrumentedPool extends ThreadPoolExecutor {
    private final ThreadLocal<Long> start = new ThreadLocal<>();
    private final Timer runTime;          // Micrometer-style metrics
    private final Counter failures;

    public InstrumentedPool(String name, int threads, int queueCap,
                            Timer runTime, Counter failures) {
        super(threads, threads, 60, TimeUnit.SECONDS,
              new ArrayBlockingQueue<>(queueCap),
              Thread.ofPlatform().name(name + "-", 0).factory(),
              new ThreadPoolExecutor.AbortPolicy());
        this.runTime = runTime;
        this.failures = failures;
    }

    @Override protected void beforeExecute(Thread t, Runnable r) {
        start.set(System.nanoTime());
    }

    @Override protected void afterExecute(Runnable r, Throwable t) {
        runTime.record(System.nanoTime() - start.get(), TimeUnit.NANOSECONDS);
        if (t == null && r instanceof Future<?> f && f.isDone() && !f.isCancelled()) {
            try { f.get(); }
            catch (ExecutionException e) { t = e.getCause(); }
            catch (InterruptedException e) { Thread.currentThread().interrupt(); }
        }
        if (t != null) failures.increment();
    }

    @Override protected void terminated() {
        System.getLogger("pool").log(System.Logger.Level.INFO, "pool terminated");
    }
}

Equal core and maximum sizes with a bounded queue make excess load fail fast instead of queueing without limit. afterExecute runs on the same worker thread, so the ThreadLocal pairing is safe. To measure queue wait too, override newTaskFor to stamp the enqueue time.

Failure modes

  • Silent task death. submit captures the exception and nobody calls get. Use execute for fire-and-forget work, or inspect futures in afterExecute as above.
  • Stopped schedules. A periodic task throws once and never runs again. Catch Throwable inside every periodic task body and log it.
  • Shutdown that never ends. Tasks block in I/O that ignores interruption, or loop without checking the flag. Check Thread.interrupted() in long loops and prefer timed or interruptible blocking calls.
  • Thread leaks. Pools created per request and never shut down keep their core threads, and non-daemon threads keep the JVM alive. Create pools at startup and give each one an owner.
  • Worker churn. Tasks that throw under execute kill and replace a worker each time, visible as ever-rising thread name suffixes.
  • close blocking forever. try-with-resources around a service-lifetime pool, or one with a stuck task, blocks the closing thread indefinitely.
  • Mis-sized scheduled pool. Raising maximumPoolSize to fix slow scheduled tasks does nothing; raise corePoolSize or move slow work onto its own pool.

Trade-offs

ThreadPoolExecutor gives fine control and observable state, at the cost of choosing queue, sizes and rejection policy yourself. ForkJoinPool trades that control for work stealing, which suits divide-and-conquer and many small tasks; see ForkJoinPool internals. Virtual-thread-per-task executors remove sizing for blocking I/O work but do not bound concurrency, so limits move to semaphores and connection pools; see virtual threads in production. Hooks such as beforeExecute and newTaskFor are powerful, but every subclass is code that must stay correct across JDK upgrades, so keep extensions thin and well tested.

What to do next

  1. List every executor in your service, who creates it and who shuts it down.
  2. Replace factory-built pools that hide unbounded queues with explicit ThreadPoolExecutor construction and named threads.
  3. Wrap every periodic task body in a catch-all that logs.
  4. Add the afterExecute check so exceptions from submitted tasks are counted.
  5. Use the shutdown, awaitTermination, shutdownNow sequence in your shutdown hook, with bounded waits.
  6. Reserve try-with-resources close for scoped executors, not service-lifetime pools.
  7. Export pool size, active count, queue depth and completed tasks, and alert on sustained queue growth.
Key takeaway: The Executor framework is layered: AbstractExecutorService turns execute into submit through FutureTask, ThreadPoolExecutor packs its lifecycle and worker count into one atomic int, workers double as non-reentrant locks so idle ones can be interrupted safely, and the scheduled pool is a heap-ordered queue on a fixed set of core threads. Those internals explain vanished exceptions, stopped schedules and hung shutdowns, and the protected hooks let you instrument all three.