Java virtual threads, final since JDK 21 (JEP 444), are threads implemented by the JVM rather than the operating system. They let you write ordinary blocking code, one thread per request, and still run hundreds of thousands of concurrent tasks, because a virtual thread that blocks on I/O gives its operating-system thread back to the scheduler instead of holding it. Before them, high concurrency meant either a large and expensive pool of platform threads or an asynchronous style built on callbacks and reactive streams.

This article explains how they work underneath: continuations, carrier threads, mounting and unmounting, which operations release the carrier and which pin it, what changed in JDK 24 and 25, and how to size and operate a service built on them.

Virtual threads: many continuations, a few carrier threadsApplication codeone virtual thread per task: Executors.newVirtualThreadPerTaskExecutor()VT 1: runningmounted on carrierVT 2: runnablequeued for a carrierVT 3: parkedwaiting on socketVT 4..N: parkedlocks, sleep, queuesHeap: stack chunksframes copied out on unmountunmountScheduler: ForkJoinPool (FIFO)parallelism = CPU cores by defaultmountCarrier 1OS threadCarrier 2OS threadCarrier NOS threadPollerI/O readiness unparks the virtual threadStill pins the carrier: native or foreign frames
Virtual threads run on a small pool of carrier threads; a blocked virtual thread unmounts, its frames move to heap stack chunks, and the poller makes it runnable again when I/O is ready.

Why thread-per-request ran out of threads

A thread-per-request server holds one thread for the whole life of a request, including the time it spends waiting on a database or another service. By Little's law, concurrency equals throughput times latency: 2,000 requests per second at 150 ms each means about 300 requests in flight, and 20,000 per second at the same latency means 3,000. Platform threads are one-to-one with OS threads. Each reserves a stack (1 MB by default on 64-bit Linux, set by -Xss; reserved address space, committed as it is touched), costs a kernel scheduling entity, and pays a kernel context switch to block and wake. Thousands work; hundreds of thousands do not.

Asynchronous code avoided the problem by never blocking, at the price of stack traces that no longer show the logical request, harder debugging and a split between blocking and non-blocking libraries. Virtual threads keep the simple model and make waiting cheap.

Why thread-per-request ran out of threads

A thread-per-request server holds one thread for the whole life of a request, including the time it spends waiting on a database or another service. By Little's law, concurrency equals throughput times latency: 2,000 requests per second at 150 ms each means about 300 requests in flight, and 20,000 per second at the same latency means 3,000. Platform threads are one-to-one with OS threads. Each reserves a stack (1 MB by default on 64-bit Linux, set by -Xss; reserved address space, committed as it is touched), costs a kernel scheduling entity, and pays a kernel context switch to block and wake. Thousands work; hundreds of thousands do not.

Asynchronous code avoided the problem by never blocking, at the price of stack traces that no longer show the logical request, harder debugging and a split between blocking and non-blocking libraries. Virtual threads keep the simple model and make waiting cheap.

How a virtual thread runs

A virtual thread is a java.lang.Thread object plus a continuation: the JVM's internal ability to suspend a running computation and resume it later. It runs by being mounted on a carrier, an ordinary platform thread owned by the scheduler. The default scheduler is a FIFO-mode ForkJoinPool with parallelism equal to the number of available processors.

When the virtual thread reaches a blocking operation that the JDK has adapted, it parks: its stack frames are copied from the carrier's stack into stack-chunk objects on the heap, and the carrier is free to mount another virtual thread. When the operation can proceed, for example when the poller sees the socket is readable, the virtual thread is unparked and queued, and some carrier, possibly a different one, mounts it and copies frames back lazily. A virtual thread's memory is therefore its small Thread object plus however many frames are live, which grows with call depth; a shallow waiting thread costs on the order of hundreds of bytes to a few kilobytes, a deep framework stack more.

Not every blocking call can unmount. Most file system operations block in the OS, so the JDK compensates by temporarily adding a carrier to the pool while the call is in progress. Code running inside a native method or a foreign function call pins the carrier: the virtual thread cannot unmount and the carrier is held until the call returns.

Platform and virtual threads side by side

PropertyPlatform threadVirtual thread
Backed byone OS thread for its whole lifeany carrier while mounted, none while parked
Stackfixed reservation, -Xssheap stack chunks, grows with depth
Creation costsystem call, kernel structuresan object allocation
Blocking I/OOS thread sleepsunmounts; carrier runs others
SchedulingOS, preemptive time slicesJVM ForkJoinPool, no time slicing
Pool it?yes, they are expensiveno, create one per task
Daemon / priorityconfigurablealways daemon, fixed priority

Both are java.lang.Thread: interruption, ThreadLocal, debuggers, stack traces and Thread.currentThread() behave the same, which is why most existing blocking code runs unchanged.

Pinning: JDK 21 versus JDK 24 and later

In JDK 21 to 23, blocking while holding a monitor, that is inside a synchronized block or method or in Object.wait(), also pinned the carrier. Enough pinned threads could exhaust the carriers, and if they waited on something only an unmounted thread could deliver, the application stalled. The standard advice was to replace synchronized around blocking I/O with ReentrantLock.

JDK 24 removed that restriction with JEP 491: virtual threads can now acquire, hold and release monitors independently of their carrier, so synchronized and Object.wait() no longer pin. That makes JDK 25, the LTS release that includes it, the practical baseline for virtual-thread services. On 21 the old advice still applies, especially to drivers and libraries you do not control. On any version, native frames still pin, and a lock is still a lock: a hot synchronized section serialises work regardless of thread type.

Using them in code

Using them is mostly a matter of not pooling them. A virtual thread is cheap to create and meant to live for one task:

try (ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor()) {
    for (Request req : requests) {
        exec.submit(() -> handle(req));        // one new virtual thread per task
    }
}                                              // close() waits for submitted tasks

Thread t = Thread.ofVirtual().name("importer-", 0).start(this::runImport);

Even the small HTTP server built into the JDK becomes a thread-per-request server with one line, because it accepts any Executor for dispatching exchanges:

HttpServer server = HttpServer.create(new InetSocketAddress(8080), 0);
server.createContext("/quote", exchange -> {
    byte[] body = quoteJson(exchange).getBytes(StandardCharsets.UTF_8);  // blocking calls are fine here
    exchange.sendResponseHeaders(200, body.length);
    try (OutputStream out = exchange.getResponseBody()) { out.write(body); }
});
server.setExecutor(Executors.newVirtualThreadPerTaskExecutor());
server.start();

Frameworks expose a switch rather than code changes: Spring Boot 3.2 and later run request handling on virtual threads with spring.threads.virtual.enabled=true, and Helidon, Quarkus and Jetty have their own virtual-thread modes. Because the thread count is no longer the limit, the limit must be expressed where it belongs, at the scarce resource:

private final Semaphore dbPermits = new Semaphore(20);   // match the connection pool size

Order load(long id) throws InterruptedException {
    dbPermits.acquire();
    try {
        return orderDao.find(id);            // blocking JDBC: the virtual thread parks
    } finally {
        dbPermits.release();
    }
}

Migrating an existing service

Migrating an existing service is mostly removal. Replace fixed thread pools used for request handling or blocking calls with a virtual-thread-per-task executor, and delete the tuning that sized them. Keep pools that exist to limit CPU work. Then go looking for implicit assumptions that there are few threads.

  • Per-thread caches in ThreadLocal that were cheap with 200 threads become per-task allocations with no reuse.
  • Thread-pool queue length used as a load-shedding signal disappears; replace it with a semaphore and a fast rejection when permits are unavailable.
  • Timeouts matter more, because nothing caps how many requests wait on a slow dependency; set connect and read timeouts on every client.
  • Libraries that call native code or hold synchronized across I/O on JDK 21 show up in JFR pinning events within minutes of a load test.

Do it behind a flag, run the old and new executors side by side under the same load, and compare latency percentiles, heap and error rates before switching over.

Structured concurrency and scoped values

Fan-out within a request is where structured concurrency fits: subtasks run in their own virtual threads, their lifetime is bounded by a scope, and failure of one cancels the others. The API is still a preview (JEP 505 in JDK 25, JEP 525 in JDK 26) and needs --enable-preview; its shape changed between previews, so expect small edits when upgrading. In the JDK 25 form:

Quote quote(long customerId) throws InterruptedException {
    try (var scope = StructuredTaskScope.open()) {          // default: all must succeed
        var profile = scope.fork(() -> profiles.get(customerId));
        var prices  = scope.fork(() -> pricing.current());
        var stock   = scope.fork(() -> inventory.levels());
        scope.join();                                       // throws if any subtask failed
        return Quote.of(profile.get(), prices.get(), stock.get());
    }
}

Scoped values, final in JDK 25 (JEP 506), carry immutable per-request context such as a tenant or trace ID into those subtasks without the mutability and per-thread copies of ThreadLocal. See structured concurrency and scoped values for the full APIs.

Worked example: sizing a quote service

Take a quote service on an 8-core machine. Each request makes three downstream HTTP calls in parallel (about 80 ms) and one database query (about 15 ms), and the target is 4,000 requests per second. In-flight requests are about 4,000 times 0.1 s, so 400, plus their subtasks, about 1,600 virtual threads. With a platform pool you would size 400 or more threads and tune them; with virtual threads you create one per request and let them park.

The database needs only 4,000 times 0.015 s, about 60 concurrent queries, so the pool is set to 64 connections and the semaphore to the same number. The carriers stay at 8 because the CPU work per request is small; if CPU time per request were 4 ms, 4,000 requests per second would need 16 cores no matter how threads are scheduled. Virtual threads raise concurrency, not compute throughput.

Failure modes

  • Pooling virtual threads in a fixed executor reintroduces the cap they were meant to remove.
  • Unbounded downstream load: 100,000 virtual threads will happily open 100,000 requests to a service sized for 500. Bound with semaphores, connection pools or rate limiters.
  • Pinning on JDK 21 through synchronized around blocking calls, often inside older JDBC drivers or caches. Upgrade, or replace with ReentrantLock.
  • Heavy ThreadLocal caches such as per-thread buffers or formatters multiply by the number of threads and are never reused; switch to shared or scoped objects.
  • CPU-bound work gains nothing and can starve other virtual threads of carriers, since scheduling is not time-sliced; run it on a bounded platform pool.
  • Native calls that block pin; keep them short or move them to platform threads.

Operating virtual-thread services

Use JDK Flight Recorder: the jdk.VirtualThreadPinned event, enabled by default with a 20 ms threshold, records where a virtual thread blocked while pinned, with a stack trace, and jdk.VirtualThreadSubmitFailed flags scheduling failures. Classic thread dumps do not scale to hundreds of thousands of threads; use jcmd <pid> Thread.dump_to_file -format=json dump.json, which groups virtual threads by their executor or structured scope. The scheduler can be tuned with jdk.virtualThreadScheduler.parallelism and jdk.virtualThreadScheduler.maxPoolSize, but the defaults suit most services. Watch heap after migration, since parked stacks now live there, and watch connection-pool wait time, which becomes the real queue. Background on the carrier pool is in ForkJoinPool and on locks in ReentrantLock.

Trade-offs

Virtual threads are the better default for I/O-bound request handling: simpler code, real stack traces and debuggers that work. Reactive frameworks still offer explicit backpressure across streams, operators for event processing and mature tooling, and remain reasonable where those features are the point. Virtual threads offer nothing for CPU-bound work, require attention to pinning on JDK 21, and move the scaling limit from threads to whatever resource is actually scarce, which you now have to bound yourself.

What to do next

  1. Move to JDK 25 if you can, so synchronized no longer pins.
  2. Switch one I/O-heavy service to a virtual-thread-per-task executor or your framework's flag.
  3. Replace thread-pool sizing with semaphores or pool limits on each downstream resource.
  4. Record JFR under load and check jdk.VirtualThreadPinned events.
  5. Audit ThreadLocal usage and move request context to scoped values.
  6. Compare p99 latency, heap and throughput before and after, and keep CPU-bound work on platform threads.
Key takeaway: Virtual threads make blocking cheap by parking continuations on the heap and freeing carrier threads, so thread-per-request scales again. Run them unpooled, bound scarce resources explicitly, use JDK 25 so monitors no longer pin, and verify with JFR.