Switching a Java service to virtual threads is usually a one-line change. Running it safely is not. The change removes the thread pool that used to cap concurrency, so load that used to queue politely at the front door now reaches your database pool, your downstream APIs and your heap. The service often gets faster, and when it fails, it fails somewhere new, with telemetry that was designed for the old shape.

This page is the runbook. It assumes you know what a virtual thread is (see JVM virtual threads architecture for the runtime and Project Loom for history and a basic migration). It covers which JDK to run, the telemetry to wire up before flipping the switch, a canary plan with explicit go and no-go criteria, six incident patterns teams actually hit, and a worked incident diagnosed from a thread dump.

Advertisement

What changes in production

With a fixed pool of 200 platform threads, the 201st concurrent request waits for a thread. That wait was a crude but effective admission control, and every dashboard you own probably watches it: active threads, pool queue length, rejected executions. With a thread per request, there is no such wait. Requests run immediately and queue instead on whatever is scarce behind them: database connections, permits in front of a dependency, locks, or CPU.

So the migration is mostly a telemetry and limits project. You must measure waiting where it now happens and put explicit limits where the thread pool used to provide them by accident.

Where requests wait, before and after virtual threadsPlatform-thread poolLoad balancerAccept queuewaits for a thread200 threadsthe visible limitDB pool, APIsrarely saturatedVirtual threadsLoad balancerThread per requestno thread limitPool and permit waitsthe new queueDB pool, APIsnow the limitTelemetry must move with the queuepool wait time, permit wait, pinned events, parked threadsThe queue does not disappear; it moves from the thread pool to the resources behind it
With platform threads the queue sits at the thread pool. With virtual threads it moves to connection pools, permits and downstream services, and dashboards must move with it.

The version matrix

JDKStatusWhat it means for production
21 (LTS)Virtual threads final (JEP 444)Blocking inside synchronized pins the carrier; long synchronized I/O can starve the scheduler
24JEP 491: synchronized no longer pins in generalRemaining pinning is mostly native frames and class loading or initialisation; jdk.tracePinnedThreads is removed
25 (LTS)Includes JEP 491The recommended baseline: pinning rarely matters and the scheduler MXBean is available for monitoring

If you are on 21 and cannot upgrade soon, audit libraries for I/O under synchronized, for example older JDBC drivers and connection pools, and prefer versions that switched to ReentrantLock. On 24 or later, JEP 491 states that a virtual thread still cannot unmount while loading a class, inside a class initializer, or while waiting for another thread to initialise a class, and JFR still records jdk.VirtualThreadPinned when native code calls back into Java that blocks.

Advertisement

Wire up telemetry before the switch

Add these signals while the service still runs on platform threads, so you have a baseline. First, stream pinning and submit-failure events from JFR inside the process and export them as metrics. The RecordingStream API lets you do this without writing recording files.

import java.time.Duration;
import jdk.jfr.consumer.RecordingStream;

public final class LoomJfrMetrics {
    public static RecordingStream start(Metrics metrics) {
        var rs = new RecordingStream();
        rs.enable("jdk.VirtualThreadPinned").withThreshold(Duration.ofMillis(20)).withStackTrace();
        rs.enable("jdk.VirtualThreadSubmitFailed").withStackTrace();
        rs.onEvent("jdk.VirtualThreadPinned", e -> {
            metrics.timer("vt.pinned").record(e.getDuration());
            metrics.counter("vt.pinned.by_method", "method", firstAppFrame(e)).increment();
        });
        rs.onEvent("jdk.VirtualThreadSubmitFailed", e -> metrics.counter("vt.submit_failed").increment());
        rs.startAsync();
        return rs;
    }

    // The top frames of a pinned event are JDK park machinery; report the first application frame.
    private static String firstAppFrame(jdk.jfr.consumer.RecordedEvent e) {
        var st = e.getStackTrace();
        if (st == null) return "?";
        for (var f : st.getFrames()) {
            var type = f.getMethod().getType().getName();
            if (!type.startsWith("java.") && !type.startsWith("jdk.") && !type.startsWith("sun.")) {
                return type + "." + f.getMethod().getName();
            }
        }
        return "jdk-only";
    }
}

Second, export the scheduler's state. In JDK 25 the jdk.management module provides VirtualThreadSchedulerMXBean, obtained with ManagementFactory.getPlatformMXBean(VirtualThreadSchedulerMXBean.class). Its getters report target parallelism, the scheduler's pool size, an estimate of mounted virtual threads, and an estimate of queued virtual threads. Queued virtual threads that stay high while CPU is not saturated is the signature of pinning or carrier starvation.

Third, make thread dumps work for virtual threads. A classic jstack output does not list them; the JSON dump does. Script it so on-call engineers can take one without remembering flags:

jcmd <pid> Thread.dump_to_file -format=json /tmp/threads-$(date +%s).json

Finally, add wait-time metrics where the queue moves to: connection acquisition time from your pool (HikariCP and others expose it), permit wait time for each semaphore in front of a dependency, and per-dependency in-flight counts. These are the dashboards that will matter after the switch.

A canary plan

  1. Put the switch behind configuration. In Spring Boot 3.2 and later it is spring.threads.virtual.enabled=true; other servers have equivalents. Replace application-owned newFixedThreadPool executors for blocking work with Executors.newVirtualThreadPerTaskExecutor() behind the same flag.
  2. Before enabling, put explicit limits in front of every dependency: database pool size with a connection timeout, a semaphore or bulkhead per downstream API sized from its real capacity, and timeouts on every call.
  3. Load test both modes with the same traffic shape, including a dependency slowed by ten times. The slow-dependency test is where virtual threads behave most differently.
  4. Canary one instance at a few percent of traffic for at least a full daily peak.
  5. Compare against control: p50 and p99 latency, error rate, CPU, heap after GC, connection wait time, pinned-event rate, queued virtual threads, and downstream error rates you can observe.
  6. Go if latency and errors are equal or better and no wait metric is growing without bound. No-go on any sustained growth in connection wait, heap after GC or downstream errors, even if your own latency looks fine; you may be exporting the overload.

Executor hygiene that pays off during incidents

A few coding conventions make the incidents below much faster to diagnose. Name your virtual threads by purpose, because a thread dump of 50,000 unnamed threads is only useful if stacks are distinctive. Never pool virtual threads: they are cheap to create, and a pool reintroduces the fixed limit you removed, this time without anyone noticing it is there. Keep a small, separate platform-thread executor for work that should not run on carriers, such as CPU-heavy transformations or calls into native libraries that block, and make that separation visible in code.

import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;

public final class Executors2 {
    // One virtual thread per task, named so dumps group by purpose.
    public static ExecutorService ioTasks(String purpose) {
        var factory = Thread.ofVirtual().name(purpose + "-", 0).factory();
        return Executors.newThreadPerTaskExecutor(factory);
    }

    // Bounded platform threads for CPU-heavy or native-blocking work, kept off the carriers.
    public static final ExecutorService CPU_AND_NATIVE = Executors.newFixedThreadPool(
            Math.max(2, Runtime.getRuntime().availableProcessors() / 2),
            Thread.ofPlatform().name("cpu-native-", 0).daemon(true).factory());
}

Leave the scheduler's system properties, jdk.virtualThreadScheduler.parallelism and jdk.virtualThreadScheduler.maxPoolSize, at their defaults unless a measurement says otherwise. Raising parallelism to compensate for starved carriers hides pinning or CPU-bound work rather than fixing it, and the extra carriers compete with the garbage collector and JIT for the same cores. Finally, check that logging context such as a request id still follows the request: frameworks that propagate it through ThreadLocal usually work with a thread per request, but hand-written executor wrappers that copy context into pooled threads need review.

Six incidents and their fixes

IncidentSymptomDiagnosisFix
Connection pool exhaustionSpike of connection timeout errors; DB fineThread dump shows thousands parked in the pool's getConnectionKeep the pool sized to the database, add a permit in front of DB-heavy endpoints, shed load early
Downstream overloadYour latency fine, partner's errors rise, then yoursIn-flight calls per dependency far above the old thread countSemaphore per dependency sized from its capacity, timeouts, retry budgets
Carrier starvationLatency climbs while CPU is moderate; queued virtual threads highPinned events with long durations, or CPU-bound loops on virtual threadsUpgrade past JDK 24, move native blocking and long CPU work to a platform-thread pool
ThreadLocal memory growthHeap grows with concurrency; large retained ThreadLocal mapsPer-thread caches such as formatters or buffers now created per requestReplace with shared thread-safe objects or pools; use ScopedValue for context
Heap from parked threadsHeap grows under slow dependencies, then GC thrashHundreds of thousands of virtual threads parked with deep stacksBound fan-out and admission; reject early instead of parking forever
Synchronized pinning on 21Tail latency under load, especially with old driversjdk.VirtualThreadPinned events inside synchronized I/OUpgrade the JDK, or upgrade the library, or swap to ReentrantLock in your own code

Two of these deserve emphasis. Downstream overload is the incident your own dashboards will miss, because your service looks healthy while it multiplies load on a partner; watch their error rate and your in-flight counts. And ThreadLocal growth is easy to cause by accident: libraries that cached expensive objects per thread assumed a few hundred long-lived threads, not one per request. ThreadLocal versus ScopedValue covers the migration for context; ScopedValue became a final API in JDK 25.

A worked incident

A checkout service moved to virtual threads on JDK 21. Two weeks later, during a sale, p99 latency rose from 300 ms to 4 seconds while CPU stayed at 40 percent and the database reported low load. The connection pool had 50 connections and its acquisition time was normal. The on-call engineer took a JSON thread dump and grouped virtual threads by the top frames of their stacks.

import json, collections, sys

def walk(containers):
    for c in containers:
        yield from c.get("threads", [])

dump = json.load(open(sys.argv[1]))["threadDump"]
groups = collections.Counter()
for t in walk(dump["threadContainers"]):
    frames = t.get("stack", [])[:4]
    groups[" <- ".join(frames)] += 1

for stack, n in groups.most_common(8):
    print(f"{n:7d}  {stack}")

The output had two striking groups. Eight threads, exactly the number of carriers on the eight-core host, were in a pricing client: one inside a synchronized method, blocked on a socket read to a slow pricing service, and seven blocked trying to enter the same monitor. About 18,000 more had not progressed past the start of the request path. On JDK 21 a virtual thread pins its carrier both while blocked inside synchronized and while blocked waiting to enter it, so those eight pricing calls held all eight carriers. Every other request, including ones that never touched pricing, sat in the scheduler queue waiting for a carrier, which the queued-thread count from the scheduler MXBean would have shown directly. The JFR stream confirmed it: pinned events with multi-second durations, all in the pricing client.

The short-term fix was to run pricing calls on a small platform-thread executor, which isolated them from the carriers. The lasting fixes were upgrading the client to a version without I/O under synchronized and planning a JDK 25 upgrade. The review also added an alert on queued virtual threads and pinned-event duration, which would have pointed at the problem within minutes. One more lesson came out of the review: the load test before rollout had used a fast pricing stub, so the slow-dependency case that triggered the incident was never exercised. The team added a fault-injection step that slows each dependency in turn by ten times and runs the test again, and that step is now part of the canary checklist. The dump format above reflects JDK 21 to 25 output; check the field names on your JDK before relying on the script.

Trade-offs and when not to switch

Virtual threads help when requests spend most of their time waiting on I/O and the thread pool was the binding limit. They do little for CPU-bound services, where the number of cores is the limit, and they add risk where you depend on libraries that block in native code. Reactive stacks that already work well have no need to convert. And a service whose real limit is a small database gains nothing from accepting more concurrent requests; it only moves the queue. Use structured concurrency for fan-out inside a request, and keep explicit limits such as a semaphore at every boundary where load leaves your process.

What to do next

  1. Pick the JDK: plan for 25, and if you must stay on 21 audit for I/O inside synchronized.
  2. Before switching, add the JFR stream, the scheduler MXBean export, a scripted JSON thread dump, connection wait time and per-dependency in-flight metrics.
  3. Put a sized limit and a timeout in front of every dependency, and size the database pool to the database, not to the old thread count.
  4. Load test with a slowed dependency in both modes, then canary for a full peak with written go and no-go criteria.
  5. Search your code and libraries for per-thread caches and replace them; move request context to ScopedValue where you can.
  6. Add alerts on queued virtual threads, pinned-event duration, heap after GC and downstream error rates, and keep the thread-dump grouping script in your runbook.
Key takeaway: Virtual threads remove the thread pool as a limit, which means the queue moves to your connection pools, permits and dependencies. Run JDK 25 if you can, wire up JFR, the scheduler MXBean and JSON thread dumps before the switch, put explicit limits in front of every dependency, canary against written criteria, and diagnose incidents by grouping parked threads by stack, which turns a vague slowdown into one named lock or pool within minutes.