Virtual threads let Java code block cheaply. A blocking call on a virtual thread parks it, its stack moves to the heap, and the small pool of carrier platform threads picks up other work. That one property removes the main reason for reactive frameworks in I/O-bound services, and it is why switching an executor can multiply throughput with no other code change.

It also creates a new class of bugs: code that compiles, passes tests and behaves differently under load because it assumed something about threads that is no longer true. Our production runbook covers the operational incidents, connection pool exhaustion and downstream overload among them, plus the JDK version matrix. This article is the code-level companion. Each pitfall gets a small reproducer, a way to detect it, and a fix. The examples target JDK 21 and later, and note where JDK 24's JEP 491 changes the picture.

Advertisement

Four facts behind every pitfall

Most pitfalls follow from four facts. First, virtual threads are cheap to create, so they are meant to be created per task and discarded; they are never a scarce resource to be pooled. Second, the scheduler is a work-stealing ForkJoinPool whose parallelism defaults to the number of cores, and it does not time-slice: a virtual thread keeps its carrier until it blocks or finishes. Third, some operations cannot unmount the virtual thread and instead hold the carrier, which is called pinning. Fourth, a virtual thread is still a java.lang.Thread, but several thread properties are fixed or invisible. The diagram shows where each of those facts turns into a failure.

Virtual threads100,000 tasks, cheapScheduler queueForkJoinPool, FIFOHeapparked stack chunksCarrier 1runs, unmounts on parkCarrier 2pinned: class initCarrier 3CPU loop, no slicingCompensating carrieradded for file I/OCommon pool (not virtual)supplyAsync, parallel streamsObservability gapThreadMXBean sees carriers onlysubmitmountpark: stack to heapwork escapes hereRed carriers are lost to the scheduler; work in the purple box never ran on a virtual thread at all
Where virtual-thread code goes wrong: carriers lost to pinning or CPU-bound loops, extra carriers created to compensate for file I/O, async APIs that hand work to the common pool instead, and monitoring that only sees platform threads.

If you want the runtime mechanics, mounting, stack chunks and the scheduler, read our virtual threads architecture article first. Everything below assumes only the four facts above.

Pitfall: pooling virtual threads

The most common migration mistake is replacing a fixed pool's thread factory and keeping the pool:

// Pitfall: a pool of 200 virtual threads behaves like a pool of 200 platform threads
ExecutorService pool = Executors.newFixedThreadPool(200, Thread.ofVirtual().factory());

This compiles and works, and gains almost nothing. Concurrency is still capped at 200, tasks still queue behind each other, and because the 200 threads are long-lived, any ThreadLocal values stay attached to them across tasks, exactly as before. Teams usually kept the pool size because it doubled as a limit on load to a database or partner API. That limit was real, and it must survive the migration, but it belongs on the resource, not on the threads:

private static final Semaphore PAYMENTS_API = new Semaphore(50);   // sized from partner capacity

try (ExecutorService exec = Executors.newVirtualThreadPerTaskExecutor()) {
    for (Order o : orders) {
        exec.submit(() -> {
            PAYMENTS_API.acquire();
            try {
                return paymentsClient.charge(o);
            } finally {
                PAYMENTS_API.release();
            }
        });
    }
}   // close() waits for every submitted task to finish

A semaphore per dependency keeps the protective limit exactly where the scarce resource is, while unrelated work runs without a cap. Our semaphore article covers fairness and timeouts with tryAcquire. To detect leftover pools, search the codebase for Thread.ofVirtual().factory() passed to anything other than a thread-per-task executor.

Advertisement

Pitfall: async APIs that escape to the common pool

Some APIs never touch the executor you configured. CompletableFuture.supplyAsync(supplier) without an executor argument runs on ForkJoinPool.commonPool(), a platform-thread pool sized to the core count. Parallel streams also use the common pool. Blocking I/O inside either occupies a scarce platform thread, so a service that moved its request handling to virtual threads can still collapse when one code path does this:

// Running on a virtual thread, but the lookups run on the common pool
List<CompletableFuture<Price>> futures = skus.stream()
    .map(sku -> CompletableFuture.supplyAsync(() -> priceService.fetch(sku)))  // blocking HTTP
    .toList();

// Fix 1: pass a virtual-thread executor explicitly
ExecutorService vexec = Executors.newVirtualThreadPerTaskExecutor();
var fixed = skus.stream()
    .map(sku -> CompletableFuture.supplyAsync(() -> priceService.fetch(sku), vexec))
    .toList();

// Fix 2: on a virtual thread, plain blocking code is usually clearer
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
    List<Future<Price>> fs = skus.stream().map(s -> exec.submit(() -> priceService.fetch(s))).toList();
    for (Future<Price> f : fs) prices.add(f.get());
}

The same applies to framework annotations for asynchronous methods, scheduled task runners and HTTP client callbacks: each has its own default executor, and enabling virtual threads for request handling in a framework does not necessarily change all of them. Check each executor by name in a thread dump rather than assuming. Parallel streams are fine for CPU-bound work on collections, which is what they were designed for; just never put blocking calls inside them.

Pitfall: pinning that survives JDK 24

On JDK 21, blocking inside a synchronized block or method, or in Object.wait(), pins the virtual thread to its carrier. JEP 491, delivered in JDK 24, removed that restriction for monitors in general, which is the main reason to run 24 or later. Two kinds of pinning remain and surprise people precisely because they assume the problem is gone. A virtual thread cannot unmount while it is running a class initializer or waiting for another thread to initialise a class, and it cannot unmount while a native frame is on its stack, for example when native code calls back into Java that then blocks.

final class TaxRules {
    // Pitfall: blocking I/O in a static initializer pins every virtual thread that touches TaxRules
    static final Map<String, BigDecimal> RATES = HttpLoader.fetchRates();   // 2 s remote call
}

// Fix: initialise lazily on first use, outside class initialisation
final class TaxRules {
    private static volatile Map<String, BigDecimal> rates;
    static Map<String, BigDecimal> rates() {
        Map<String, BigDecimal> r = rates;
        if (r == null) {
            synchronized (TaxRules.class) {          // fine on JDK 24+; use a ReentrantLock on 21
                if ((r = rates) == null) rates = r = HttpLoader.fetchRates();
            }
        }
        return r;
    }
}

During the two seconds of the blocking initializer, every other virtual thread that touches TaxRules waits for the class and pins a carrier; with 8 carriers and a burst of requests, the whole service stalls. Detect pinning with the JFR event jdk.VirtualThreadPinned, which records blocked operations over a threshold with a stack trace: jcmd <pid> JFR.start settings=profile then look for the event in the recording. The old jdk.tracePinnedThreads system property was removed in JDK 24, so do not rely on it.

Pitfall: CPU-bound work and no time slicing

Because the scheduler does not preempt, a virtual thread running a long computation holds its carrier until it finishes. Put a few image resizes or a large JSON transformation on virtual threads alongside latency-sensitive request handlers and the handlers wait in the queue behind them, even though the machine is not saturated by I/O. Virtual threads make waiting cheap; they do not make computing faster, and they give no fairness for CPU work.

// Keep CPU-heavy work on a bounded platform pool sized to the cores you can spare
private static final ExecutorService CPU =
    Executors.newFixedThreadPool(Math.max(1, Runtime.getRuntime().availableProcessors() - 2));

Thumbnail handle(Request r) throws Exception {
    byte[] original = blobStore.get(r.key());              // blocking I/O, virtual thread parks
    return CPU.submit(() -> Resizer.resize(original, 256)) // CPU work on a platform thread
              .get();                                       // virtual thread parks while waiting
}

The handler stays on a virtual thread and parks cheaply while the platform pool does the computation, so the I/O path keeps its carriers. A tempting alternative, sprinkling Thread.yield() through loops, is fragile; isolate the work instead.

Pitfall: tools that cannot see virtual threads

Tools written for platform threads quietly miss virtual threads. Thread.getAllStackTraces() returns only platform threads, and the ThreadMXBean thread counts and dumps report platform threads too, so a dashboard of thread count stays flat while a hundred thousand virtual threads are parked. The traditional jcmd <pid> Thread.print dump also lists platform threads. Use the newer dump, which includes virtual threads and can write JSON for tooling:

jcmd <pid> Thread.dump_to_file -format=json /tmp/threads.json

Virtual threads also have empty names by default, which makes any dump hard to read. Name them through the builder so dumps and logs show what each thread is for:

ThreadFactory f = Thread.ofVirtual().name("order-worker-", 0).factory();   // order-worker-0, -1, ...
ExecutorService exec = Executors.newThreadPerTaskExecutor(f);

Replace thread-count gauges with counts you control: in-flight requests, semaphore permits in use, and executor submissions. On JDK 25 the scheduler MXBean adds carrier-level metrics; the production runbook shows how to export them.

Pitfall: assumptions about thread identity

Several thread properties are fixed for virtual threads. They are always daemon threads, and setDaemon(false) throws IllegalArgumentException. Their priority is always normal, and setPriority has no effect. The daemon rule causes a real bug in command-line tools and batch jobs:

public static void main(String[] args) {
    Thread.startVirtualThread(() -> exportAllInvoices());   // Pitfall: JVM may exit immediately
}                                                           // daemon threads do not keep it alive

public static void main(String[] args) throws Exception {
    Thread t = Thread.startVirtualThread(() -> exportAllInvoices());
    t.join();                                               // Fix: wait explicitly, or use an executor
}

Code that used priorities to favour some work silently loses that behaviour; replace it with separate executors or admission control. Per-thread caches in ThreadLocal, once amortised over a few hundred pooled threads, now run once per task, so an expensive formatter or buffer is rebuilt for every request. Our ThreadLocal versus ScopedValue article covers the replacements.

Pitfall: file I/O and carrier compensation

Most file system operations cannot be made non-blocking on every operating system, so a virtual thread reading a file holds its carrier. The JDK compensates by temporarily adding carriers to the scheduler, up to jdk.virtualThreadScheduler.maxPoolSize. That keeps the system live, but a job that reads thousands of files from virtual threads at once can drive the carrier count far above the core count, with the extra platform threads, memory and context switching that implies. Bound file I/O concurrency with a semaphore, the same way as any other scarce resource. A permit count of a few times the device queue depth is a reasonable start for local disks; network file systems need their own measurement.

Pitfall: executor lifecycle and lost exceptions

Two lifecycle details catch people. ExecutorService is AutoCloseable since JDK 19, and close() waits for all submitted tasks; a try-with-resources block around a per-task executor is therefore a join point, and a request handler that wraps a fan-out this way will wait for the slowest subtask unless each subtask has its own timeout. And submit() captures exceptions in the returned future: if nobody calls get(), failures disappear. Use invokeAll or collect the futures and check them. StructuredTaskScope packages this pattern, cancelling siblings when one fails, but it is still a preview API through JDK 25, so check its status on your JDK; see our structured concurrency article.

Worked example: an enrichment job that got slower

Consider an enrichment job on JDK 21 that read 40,000 customer records, called a profile API and a fraud API for each, and wrote results to a database. It moved from a 64-thread pool to virtual threads and got slower under load. Four pitfalls from this article explain it, and they were fixed in this order:

  1. The pool had been kept as newFixedThreadPool(64, Thread.ofVirtual().factory()). Replacing it with a per-task executor plus semaphores of 100 for the profile API and 40 for the fraud API raised concurrency where it was allowed and protected the fraud service.
  2. Fraud scores were fetched with supplyAsync and no executor, so they ran on the common pool's 7 threads. Passing the virtual executor removed the bottleneck.
  3. The fraud client library did its HTTP calls inside a synchronized method, so JFR showed jdk.VirtualThreadPinned events of 300 ms. Upgrading to JDK 25 removed them, with no code change.
  4. The job's main started a coordinating virtual thread and returned, so batch runs sometimes ended early with no error. Joining the coordinator fixed it.

Throughput went from roughly 1,100 to 6,500 records per minute in this illustrative run, limited by the fraud API's permit count, which is the limit you want to be bound by.

Trade-offs

Virtual threads are the right default for request-per-thread services and fan-out over I/O. They are the wrong tool for CPU-bound pipelines, for code that depends on thread priority, and for hot paths through native libraries that block. A mixed design, virtual threads for I/O with a bounded platform pool for computation and explicit semaphores for every external dependency, gets the benefits without the pitfalls. Expect to spend the migration effort on limits and observability, not on the executor change itself.

What to do next

  1. Search for Thread.ofVirtual().factory() passed to fixed pools and replace them with per-task executors plus a semaphore per dependency.
  2. Find every supplyAsync, runAsync and parallelStream on an I/O path and pass an explicit executor or rewrite as plain blocking code.
  3. Enable the jdk.VirtualThreadPinned JFR event in a load test, and audit static initializers for I/O.
  4. Move CPU-heavy steps to a bounded platform pool and check request latency under a mixed load.
  5. Switch thread dumps to jcmd Thread.dump_to_file -format=json and name virtual threads through the builder.
  6. Check CLI tools and batch jobs for virtual threads started from main without a join.
  7. Plan the move to JDK 24 or later, preferably the JDK 25 LTS, so monitor pinning stops being a concern.
Key takeaway: Virtual threads make blocking cheap, but they break code that assumed threads were scarce, preemptive or visible. Do not pool them; put limits on resources with semaphores. Pass explicit executors to CompletableFuture, keep blocking out of parallel streams, keep I/O out of static initializers, move CPU work to a bounded platform pool, and use the new thread dump format. Remember that virtual threads are always daemons, and plan the move to JDK 24 or later.