Since JDK 21, Java has two kinds of thread behind one class. A platform thread is the thread Java always had: a thin wrapper around an operating-system thread. A virtual thread is a java.lang.Thread that the JVM schedules itself, running it on a small pool of platform threads and taking it off whenever it blocks. Both run the same code, use the same locks and show up in the same stack traces, so it is tempting to treat the choice as a flag. It is not. The two models differ in what a thread costs, who decides when it runs, what happens when it blocks and what your tools can see.

This article puts the two architectures side by side: cost, blocking, sizing with Little's law, migration and failure. The internals of virtual threads are covered in how virtual threads work; this page is about the comparison and the decision.

Two thread models behind one class

A platform thread maps one to one onto a kernel thread. When you call new Thread(r).start(), the JVM asks the operating system for a thread, which gets its own native stack, kernel bookkeeping and a slot in the kernel run queue. The kernel decides when it runs, preempts it when its time slice ends, and parks it in the kernel when it blocks on a socket or a lock. The JVM is mostly a passenger.

A virtual thread maps many to few. It is a Java object holding a continuation, the captured state of a computation that can be suspended and resumed. To run, it is mounted on a carrier, an ordinary platform thread owned by the JVM's scheduler, a ForkJoinPool in FIFO mode whose parallelism defaults to the number of available processors. When the code blocks on something the JDK knows how to wait for asynchronously, such as a socket read or a ReentrantLock, the JVM unmounts it: its frames are copied to the heap, the carrier goes off to run another virtual thread, and when the I/O or lock is ready the virtual thread is rescheduled, possibly on a different carrier.

Both are instances of Thread; ThreadLocal, interrupts, synchronized and java.util.concurrent all work. The differences sit underneath the API.

Platform threads (1:1) versus virtual threads (M:N)Platform threadsThread AThread BThread COS thread1 MB stackOS thread1 MB stackOS thread1 MB stackKernel schedulerpreemptive, time-slicedVirtual threadsVT 1VT 2VT 3VT ...Parked stacks live on the heapstack chunks, sized by real depthJVM scheduler (ForkJoinPool, FIFO)mount on run, unmount on blockCarrierOS threadCarrierOS threadCarrierOS threadKernel schedulersees only the carriersLeft: every Java thread is a kernel thread. Right: many virtual threads share a few carriers, roughly one per core.
The kernel schedules platform threads directly. Virtual threads are scheduled by the JVM onto carrier threads, and the kernel sees only the carriers.

What each kind of thread costs

Memory. A platform thread reserves a native stack of a fixed maximum size, set by -Xss or the Thread constructor; on 64-bit Linux the default is typically 1 MB. That is a reservation of address space, and the operating system commits pages only as the stack is touched, so a shallow thread uses far less physical memory than its reservation. The real limits are usually ulimit -u, kernel.threads-max, memory-mapping counts, container PID limits and scheduler overhead, long before RAM.

A virtual thread has no fixed stack. While mounted it runs on the carrier's stack; when it unmounts, the frames it actually uses are copied into heap objects called stack chunks. Its memory is proportional to its real call depth and lives on the garbage-collected heap. A million parked virtual threads with shallow stacks fit in an ordinary heap; with deep framework stacks they will not, and the symptom is GC pressure rather than a thread-creation error.

Creation. Starting a platform thread is a system call plus stack setup, which is why server code pools them. Starting a virtual thread is allocating an object and submitting it to a queue, cheap enough that the intended pattern is one virtual thread per task, never a pool.

Switching. A platform-thread switch is a kernel context switch. A virtual-thread switch is a user-space copy of changed frames to or from the heap: no kernel, but not free, since the copy grows with stack depth. Measure both on your own hardware with the harness below rather than trusting published numbers.

Blocking and scheduling compared

A blocked platform thread costs a kernel thread for as long as it waits. A blocked virtual thread costs a heap object and frees its carrier, provided the JDK can unmount it. Where it cannot, it is pinned or captures its carrier, occupying one of your few carriers.

OperationPlatform threadVirtual thread
Socket read or write, HttpClient, JDBC over socketsKernel thread blocksUnmounts; carrier is freed
ReentrantLock, Semaphore, BlockingQueueParks in the kernelUnmounts
Thread.sleep, LockSupport.parkParks in the kernelUnmounts
synchronized and Object.waitParks in the kernelJDK 21: pins the carrier. JDK 24 and later (JEP 491): unmounts in general
File I/OKernel thread blocksCaptures the carrier; the scheduler may temporarily add a carrier to compensate
Native code and JNI frames on the stackKernel thread blocksPinned
Class loading or a class initializerNormalPinned, even on JDK 24 and later
Long CPU loopPreempted by kernel time slicingNot preempted; holds the carrier until it blocks or ends

The last row is the second big difference: scheduling policy. The kernel time-slices platform threads, so a CPU-heavy thread cannot starve its neighbours for long. The virtual-thread scheduler is cooperative at blocking points: a virtual thread runs until it blocks or finishes. Eight CPU-bound virtual threads on eight cores occupy every carrier while thousands of I/O-bound ones wait. Virtual threads also ignore setPriority and are always daemon threads. Common real-world traps are catalogued in virtual thread pitfalls.

Same code, both models: a harness

The API is deliberately the same, so one program can exercise both models. This harness starts N tasks that each block for a fixed time and reports wall time and peak live threads. Run it with platform and virtual threads at increasing N, and with the stack depth and sleep time adjusted to resemble your service.

// ThreadModels.java - run: java ThreadModels virtual 100000   or   java ThreadModels platform 2000
import java.time.Duration;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicInteger;

public class ThreadModels {
    static final AtomicInteger live = new AtomicInteger();
    static final AtomicInteger peak = new AtomicInteger();

    static void task() {
        int now = live.incrementAndGet();
        peak.accumulateAndGet(now, Math::max);
        try {
            Thread.sleep(Duration.ofMillis(200));   // stands in for a downstream call
        } catch (InterruptedException e) {
            Thread.currentThread().interrupt();
        } finally {
            live.decrementAndGet();
        }
    }

    public static void main(String[] args) {
        boolean virtual = args[0].equals("virtual");
        int n = Integer.parseInt(args[1]);
        long start = System.nanoTime();
        try (ExecutorService ex = virtual
                ? Executors.newVirtualThreadPerTaskExecutor()
                : Executors.newFixedThreadPool(200)) {     // a typical server pool
            for (int i = 0; i < n; i++) ex.submit(ThreadModels::task);
        }   // close() waits for every task
        long ms = (System.nanoTime() - start) / 1_000_000;
        System.out.println(args[0] + " n=" + n + " wall_ms=" + ms + " peak_live=" + peak.get());
    }
}

With a pool of 200, wall time grows linearly with N: 10,000 tasks take about 50 rounds of 200 ms. With virtual threads, peak concurrency equals N and wall time stays near 200 ms until heap becomes the limit. Virtual threads remove the thread limit, not the work.

Sizing with queueing arithmetic: a worked example

Whether virtual threads help is arithmetic. Little's law says the average number of requests in flight equals arrival rate times time in system: L = λ × W. In a thread-per-request server each in-flight request holds a thread, so L is the number of threads you need.

Worked example. A checkout service receives 2,000 requests per second. Each takes 150 ms, of which 10 ms is CPU and 140 ms is waiting on a payment API and a database. Threads needed: 2,000 × 0.15 = 300. CPU needed: 2,000 × 0.010 = 20 core-seconds per second, so 20 cores. A 300-thread platform pool handles this; virtual threads change little.

Now the payment API degrades and its latency rises to 2 seconds. W becomes about 2.01 s and L becomes 4,020. The 300-thread pool saturates and latency explodes for every endpoint sharing it, including ones that never call payments. With virtual threads, 4,020 threads are cheap and the service keeps accepting work, but 4,020 concurrent calls now hit a payment API that is already struggling.

So the rule is: virtual threads help when L is large because W is dominated by waiting, and CPU is not the bottleneck. They do not help when work is CPU-bound, because L is then bounded by cores, and they move rather than remove the bottleneck when a downstream has a fixed capacity. The platform pool used to act as an accidental limit on downstream load; after migration you must put that limit back explicitly.

Migrating a thread-pool service

A safe migration keeps the code and replaces implicit pool limits with explicit ones.

// Before: the pool size is both a thread budget and a hidden limit on downstream load.
ExecutorService workers = Executors.newFixedThreadPool(200);

// After: one virtual thread per task, and the limit made explicit per dependency.
ExecutorService workers = Executors.newVirtualThreadPerTaskExecutor();
Semaphore paymentPermits = new Semaphore(64);          // sized from the payment API's capacity

PaymentResult charge(Order o) throws InterruptedException {
    if (!paymentPermits.tryAcquire(50, TimeUnit.MILLISECONDS)) {
        throw new RejectedExecutionException("payment bulkhead full");   // shed load early
    }
    try {
        return paymentClient.charge(o);                // blocking call; the virtual thread unmounts
    } finally {
        paymentPermits.release();
    }
}

// CPU-heavy work stays on a bounded platform pool so it cannot monopolise carriers.
ExecutorService cpuPool = Executors.newFixedThreadPool(Runtime.getRuntime().availableProcessors());

In frameworks the switch is usually a setting: Spring Boot 3.2 and later run request handling on virtual threads with spring.threads.virtual.enabled=true, and recent Tomcat and Jetty versions offer virtual-thread executors. Underneath: never pool virtual threads, keep connection pools sized to the database, and audit ThreadLocal caches, which become per-request allocations. For request context that used ThreadLocal, ThreadLocal versus ScopedValue covers the move to ScopedValue, final since JDK 25. A rollout plan with canaries and telemetry is in virtual threads in production.

Observability differences

Tools built for platform threads see virtual threads only partially, and the difference matters during an incident.

  • Thread dumps. jstack and Thread.getAllStackTraces() show platform threads, including carriers, but not parked virtual threads. Use jcmd <pid> Thread.dump_to_file -format=json <file>, which includes virtual threads grouped by container, and group the result by stack to find the lock or pool they are waiting on.
  • Thread counts. ThreadMXBean counts platform threads. A dashboard alarm on thread count goes quiet after migration, so replace it with in-flight request counts and permits in use.
  • JFR. jdk.VirtualThreadPinned records pinning above a threshold, jdk.VirtualThreadSubmitFailed records scheduling failures, and the start and end events can be enabled for sampling. This is the main source of truth for carrier problems.

Failure modes side by side

FailurePlatform-thread formVirtual-thread form
Too much concurrencyPool saturates; requests queue; RejectedExecutionException or timeoutsNo queue in front; downstream pools and APIs overload; heap grows with parked stacks
Slow dependencyAll shared-pool endpoints stallOnly callers of that dependency stall, if each has its own bulkhead
CPU hogKernel time slicing keeps others responsiveCarriers monopolised; I/O-bound threads wait despite idle CPUs elsewhere
Blocking that cannot unmountNot applicablePinned carriers; on JDK 21 mostly synchronized I/O, on 24 and later native frames and class initialisation
Leaked threadsVisible in jstack; capped by the poolInvisible to jstack; unbounded; found only in JSON dumps or heap
Per-thread caches200 copies, fineOne copy per request, allocation and GC churn

Trade-offs and a decision matrix

Neither model is better in general; each is right for a shape of work.

ChooseWhen
Virtual threadsRequest-per-thread servers and clients whose time is mostly waiting on network I/O; fan-out to many services; code you want to keep synchronous and readable instead of rewriting into callbacks or reactive chains
Platform threadsCPU-bound computation, where cores are the limit; code that blocks in native libraries or JNI; work that needs priorities, time slicing or a dedicated thread, such as a game loop or audio thread; JDK 21 code with I/O under synchronized that you cannot change
BothMost real services: virtual threads for request handling and I/O, a small bounded platform pool for CPU-heavy steps, and explicit semaphores in front of every dependency

Compared with reactive code, virtual threads keep ordinary stack traces and exception handling but give up explicit backpressure, which you rebuild with semaphores. Compared with a tuned pool, they remove saturation failures but move overload downstream. For splitting work inside a request, structured concurrency pairs naturally with virtual threads; it remained a preview API through JDK 25, so check its status in your JDK before depending on it.

What to do next

  1. Measure L for your service with Little's law: arrival rate times average latency, split into CPU time and wait time. If L is in the hundreds and CPU-bound, stop here; a platform pool is fine.
  2. Run the harness above on your production JDK and hardware with your real stack depth, and record wall time, heap and peak concurrency for both models.
  3. Run JDK 25 if you can, so that synchronized no longer pins; on JDK 21, find I/O under synchronized with the jdk.VirtualThreadPinned JFR event.
  4. Before switching, inventory every limit the old pool enforced implicitly, and add a semaphore or bulkhead per dependency sized from that dependency's capacity.
  5. Move CPU-heavy steps to a bounded platform pool, and remove any pooling of virtual threads.
  6. Replace thread-count dashboards with in-flight requests, permits in use, JFR pinning events and JSON thread dumps grouped by stack.
  7. Canary the switch on one service, compare tail latency and downstream error rates against the platform-thread baseline, and keep the old executor behind a flag until the numbers hold under a dependency slowdown.
Key takeaway: A platform thread is an operating-system thread that the kernel schedules and time-slices; a virtual thread is a continuation the JVM mounts on a few carrier threads and unmounts whenever it blocks on I/O or a lock. Virtual threads make waiting cheap, so they help when Little's law says you need many concurrent requests that mostly wait, and they do nothing for CPU-bound work. Keep CPU-heavy and native-blocking work on bounded platform pools, put explicit limits in front of every dependency because the old pool no longer protects them, run JDK 24 or later to avoid synchronized pinning, and switch your monitoring to JFR events and JSON thread dumps.