A virtual thread is a java.lang.Thread that the JVM, not the operating system, schedules. Since JDK 21 (JEP 444) you can start millions of them, write plain blocking code in each, and let the runtime turn every blocking call into a cheap suspension. That promise is simple to state and easy to misuse, because the mechanism underneath has sharp edges: some operations cannot suspend, the scheduler never preempts, suspended stacks live on the heap, and removing the thread limit exposes whatever limit was hiding behind it.

This article traces the runtime rather than the API: one virtual thread through a blocking read and back, the cost of freezing and thawing stacks, the scheduler, how each kind of blocking is handled on JDK 21 and JDK 24+, and what to bound once threads are no longer scarce. The migration story and the wider Project Loom history are covered in Project Loom, in depth; here the goal is to let you predict what the runtime will do with your code.

Advertisement

What a virtual thread is made of

Strip away the API and a virtual thread is three things. First, a Thread object, so identity, interrupt status, thread-locals and the name all work as they always did. Second, a continuation: an internal JDK object that holds a runnable body and can be suspended at a point and resumed later from that point. Third, a reference to a scheduler, by default a dedicated ForkJoinPool whose workers are ordinary platform threads called carriers.

Running a virtual thread means mounting its continuation on a carrier: the carrier's native stack is used to execute the virtual thread's frames, and Thread.currentThread() returns the virtual thread, not the carrier. Blocking means the opposite: the continuation yields, its frames are moved off the carrier's stack, and the carrier returns to the pool to run something else. The OS never learns that the virtual thread exists. It sees a small, fixed number of carrier threads that are almost always busy.

That design explains the core economics. A platform thread reserves a native stack (commonly 1 MB of address space by default on 64-bit Linux, committed as it is touched) and costs a kernel context switch to block and resume. A parked virtual thread costs only the heap bytes of its frames plus a few hundred bytes of bookkeeping, and resuming it is a task submission and a memory copy. Nothing about this makes code run faster; it makes waiting cheap.

One virtual thread blocking on a socket read: park, freeze, unmount, poll, unpark, thawVirtual threadThread object + continuationCarrier threadForkJoinPool workerScheduler queueFIFO, work-stealingSocket readwould block: EAGAINFreezeframes copied to heapStack chunkordinary GC'd objectPoller threadepoll / kqueue / IOCPUnparkresubmit continuationThawcopy back a few framestake taskmountblocking callparkstoreregister fdfd readyenqueueon remountWhile parked the virtual thread holds no carrier and no OS thread: only its heap stack chunk and Thread object.The carrier that ran it moves straight on to the next queued task; the resume may happen on a different carrier.
The blocking path for a socket read. Only the heap stack chunk survives while the thread waits; any carrier can resume it.

The lifecycle, traced

Follow a request handler started with Thread.ofVirtual().start(task) or by a newVirtualThreadPerTaskExecutor.

  1. Start. The thread moves from new to started and its continuation is submitted to the scheduler as a task. No OS thread is created.
  2. Mount. A carrier takes the task, marks the virtual thread running, and calls into the continuation. The handler's frames now sit on the carrier's native stack.
  3. Block. The handler calls InputStream.read on a socket. JDK socket implementations used from virtual threads keep the underlying file descriptor non-blocking, so the read returns 'would block' instead of sleeping in the kernel.
  4. Park. The JDK registers the descriptor with a poller (a platform thread driving epoll, kqueue or IOCP) and calls the internal park. The state becomes parking.
  5. Freeze and unmount. The continuation yields. Its frames are copied from the carrier stack into a heap object, the stack chunk, and the carrier is released. The state becomes parked.
  6. Unpark. Data arrives; the poller sees the descriptor ready and unparks the thread, which resubmits the continuation to the scheduler.
  7. Thaw and continue. Some carrier, not necessarily the original, takes the task and copies frames back from the chunk. The read retries, succeeds and returns to the handler as if it had simply blocked.

Internally VirtualThread keeps a state field with more values than this (timed parking, pinned, yielding and, since JDK 24, states for blocking on monitors), but the simplified sequence above is what matters for reasoning. Two consequences follow. A virtual thread can change carriers at every blocking point, so never cache carrier identity or rely on per-carrier native state. And a virtual thread only gives up its carrier at a blocking point: between them it runs to completion on the carrier it has.

Advertisement

Stack chunks: where a parked stack lives

Freezing copies the frames between the continuation's entry and the yield point into a stack chunk, which is an ordinary heap object managed by the garbage collector. HotSpot keeps this cheap in two ways. Freezing copies only the frames that changed since the last thaw when it can, and thawing is lazy: it copies back only the top few frames and installs a return barrier, so that when execution returns past them the next frames are thawed on demand. A deep framework stack that parks and resumes often mostly stays on the heap until it is actually needed.

The cost that remains is heap. A worked estimate: a handler parked about 60 frames deep with typical frame sizes might hold a few kilobytes in its chunk. At 50,000 concurrently parked requests at 4 KB each, that is roughly 200 MB of live heap which the collector must trace, and which moves between generations like any long-lived object if requests wait long. That is far cheaper than 50,000 platform threads, but it shows up as old-generation growth rather than as thread count, so check a heap histogram under load.

Stack depth therefore matters more than it used to. Deep recursive frameworks, huge local arrays and many live locals all enlarge chunks. Thread-locals matter too: each virtual thread has its own map, so a library that caches a 64 KB buffer in a ThreadLocal now allocates one per request instead of one per pooled thread. ThreadLocal vs ScopedValue covers the replacement pattern.

The scheduler: FIFO, no time slices, and compensation

The default scheduler is a ForkJoinPool running in FIFO (asynchronous) mode. Its parallelism, the number of carriers normally running, is set by jdk.virtualThreadScheduler.parallelism and defaults to the number of available processors; jdk.virtualThreadScheduler.maxPoolSize caps how many carriers can exist when the pool compensates, and defaults to 256. Idle carriers steal from busy ones exactly as described in work stealing.

The scheduler does not time-slice. A virtual thread that computes for two seconds without blocking holds its carrier for two seconds. With eight carriers, eight such threads stall every other virtual thread in the process, including the ones that would answer your health check. This is the most common production surprise: CPU-heavy work (JSON of a 50 MB body, image resizing, big regex backtracking) belongs on a bounded platform-thread pool, and the virtual thread should hand it off and wait for the result, which parks cleanly.

Compensation is the pool's answer to operations that must block the carrier itself. When a virtual thread performs a blocking file read, or on JDK 21 calls Object.wait(), the JDK temporarily adds a carrier so parallelism is preserved, up to maxPoolSize. That keeps throughput up but means your carrier count can briefly exceed the core count, and a flood of slow file operations can still exhaust the cap.

How each blocking operation is handled

OperationJDK 21JDK 24 and later
Socket and pipe I/O, HttpClientParks; poller wakes itSame
ReentrantLock, BlockingQueue, Semaphore, futuresParks (LockSupport.park)Same
Thread.sleep, timed waitsTimed parkSame
Blocking inside synchronizedPins the carrierUnmounts (JEP 491)
Object.wait()Pins, pool compensatesUnmounts (JEP 491)
File I/OBlocks carrier, pool compensatesSame
Blocking in class loading or a class initializerPinsStill pins
Native code / JNI that blocksPinsStill pins

Pinning means the virtual thread blocks while mounted, so its carrier blocks with it. Short pins are harmless; pins held across slow I/O leave every carrier asleep. On JDK 21 the classic cause is a JDBC driver or cache that does network I/O inside synchronized; the usual fix was to replace the monitor with a ReentrantLock. JEP 491 (JDK 24) made the JVM track monitors held by virtual threads so that synchronized and Object.wait() unmount instead, and removed the jdk.tracePinnedThreads property. The jdk.VirtualThreadPinned JFR event remains and now reports the pin reason and the carrier, which is what you use to find the native-code and class-initialisation cases that are still left.

A worked example: fan-out with an explicit bulkhead

A catalogue endpoint prices 10,000 SKUs by calling a pricing service. With a pool of 200 platform threads the pool itself limited concurrency to 200, and the pricing service was sized for that. Switch to one virtual thread per call and the limit vanishes: 10,000 requests hit pricing at once, its connection pool saturates, latencies climb, timeouts trigger retries, and the outage is blamed on virtual threads. The fix is to make the implicit limit explicit.

import java.net.URI;
import java.net.http.*;
import java.time.Duration;
import java.util.concurrent.*;

public final class PriceFanOut {
    // The thread is no longer the scarce resource; the downstream is. Bound it explicitly.
    private static final Semaphore PRICING_PERMITS = new Semaphore(64);
    private static final HttpClient HTTP = HttpClient.newBuilder()
            .connectTimeout(Duration.ofSeconds(2)).build();

    static String price(String sku) throws Exception {
        if (!PRICING_PERMITS.tryAcquire(200, TimeUnit.MILLISECONDS)) {
            throw new RejectedExecutionException("pricing bulkhead full");   // shed, don't queue forever
        }
        try {
            var req = HttpRequest.newBuilder(URI.create("http://pricing/v1/" + sku))
                    .timeout(Duration.ofMillis(800)).build();
            return HTTP.send(req, HttpResponse.BodyHandlers.ofString()).body();  // parks, never pins
        } finally {
            PRICING_PERMITS.release();
        }
    }

    public static void main(String[] args) throws Exception {
        try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {       // one new thread per task
            var futures = new java.util.ArrayList<Future<String>>();
            for (int i = 0; i < 10_000; i++) {
                String sku = "sku-" + i;
                futures.add(exec.submit(() -> price(sku)));
            }
            for (var f : futures) {
                try { f.get(); } catch (ExecutionException e) { /* count, log, degrade */ }
            }
        }   // close() waits for every task, so no thread outlives this block
    }
}

Three details carry the design. The Semaphore is the bulkhead: at most 64 calls are in flight to pricing no matter how many virtual threads exist, and acquiring it parks rather than pins. tryAcquire with a timeout sheds load instead of letting an unbounded queue of parked threads grow on the heap. And the try-with-resources executor gives structure: close() waits for every task, so no thread outlives the request. For richer cancellation and error propagation use a scope, as described in structured concurrency architecture.

To see what the runtime is actually doing, use JFR and the JSON thread dump:

# JFR: pinning (reason + carrier on JDK 24+) and failed submissions to the scheduler
java -XX:StartFlightRecording=duration=120s,filename=vt.jfr -jar app.jar
jfr print --events jdk.VirtualThreadPinned,jdk.VirtualThreadSubmitFailed vt.jfr

# Thread dump that includes virtual threads (jstack shows platform threads only)
jcmd <pid> Thread.dump_to_file -format=json /tmp/threads.json

# Scheduler sizing: leave at defaults unless a measurement says otherwise
java -Djdk.virtualThreadScheduler.parallelism=8 \
     -Djdk.virtualThreadScheduler.maxPoolSize=256 -jar app.jar

The thread dump groups virtual threads by the executor that created them, so a pile-up of parked threads is easy to spot. Java Flight Recorder covers recording configuration and overhead.

Capacity engineering after the thread limit

Little's law says requests in flight equal arrival rate times time in system. At 2,000 requests per second and 150 ms each, you have 300 in flight on average; a slow dependency that pushes latency to 3 s raises that to 6,000. A platform-thread pool of 400 would have rejected or queued the excess, a crude but real form of back-pressure. Virtual threads accept all 6,000, so the pressure lands on whatever is actually finite.

Walk the request path and list the finite things: database connection pools, HTTP connection pools per host, downstream rate limits, file descriptors, heap for parked stacks and request buffers, and the CPU itself. Each needs an explicit bound and an explicit behaviour when full. A 20-connection database pool in front of 6,000 virtual threads is a queue of 5,980 parked threads, each holding its request buffers, all waiting on the pool's acquisition timeout. Size semaphores from the downstream's measured capacity, give every acquire a timeout, and return a fast 503 or a degraded answer rather than queueing forever.

Failure modes

  • Carrier starvation by CPU work. Symptoms: all carriers busy, latency spikes across unrelated endpoints, no pinning events. Move compute to a bounded platform pool.
  • Pinning on native or class-init paths. Symptoms: throughput collapses to the carrier count under load while CPU is idle. Find it with jdk.VirtualThreadPinned.
  • Pooling virtual threads. A fixed pool reintroduces the limit; create one per task.
  • Downstream overload. The thread limit was the only back-pressure; add bulkheads before switching.
  • Heap growth. Millions of parked threads with deep stacks or big thread-locals show up as old-generation pressure and longer GC cycles.

Trade-offs

Virtual threads give you blocking code with the concurrency of asynchronous code, plus stack traces and profilers that make sense. Reactive styles still win where back-pressure must be built into the stream itself. The costs are no preemption, residual pinning, heap-resident stacks, and the loss of the thread pool as an accidental throttle.

What to do next

  1. Run on JDK 24 or later if you can, so synchronized no longer pins; otherwise audit monitors held across I/O.
  2. Record a load test with jdk.VirtualThreadPinned enabled and fix every pin longer than a few milliseconds.
  3. List every finite downstream resource and put a Semaphore with a timeout in front of each.
  4. Move CPU-heavy work to a bounded platform pool and keep the scheduler at its defaults unless a measurement disagrees.
  5. Replace per-thread caches in ThreadLocal with shared or scoped alternatives.
  6. Switch thread observability to jcmd Thread.dump_to_file and track heap held by stack chunks.
Key takeaway: A virtual thread is a Thread object, a continuation and a scheduler reference. Blocking parks it, freezes its frames into a heap stack chunk and frees the carrier; unparking resubmits it and thaws frames lazily, possibly on another carrier. The scheduler is a FIFO ForkJoinPool sized to the core count, with no time slicing and temporary compensation for operations that block the carrier. Since JDK 24 synchronized and Object.wait unmount, but native code and class initialisation still pin. Once threads are free, bound the real resources explicitly, keep compute off the carriers, and watch the heap and JFR rather than thread counts.