Since JDK 21, Java has two kinds of thread behind one class. A platform thread is the thread Java always had: a thin wrapper around an operating-system thread. A virtual thread is a java.lang.Thread that the JVM schedules itself, running it on a small pool of platform threads and taking it off whenever it blocks. Both run the same code, use the same locks and show up in the same stack traces, so it is tempting to treat the choice as a flag. It is not. The two models differ in what a thread costs, who decides when it runs, what happens when it blocks and what your tools can see.
This article puts the two architectures side by side: cost, blocking, sizing with Little's law, migration and failure. The internals of virtual threads are covered in how virtual threads work; this page is about the comparison and the decision.
Two thread models behind one class
A platform thread maps one to one onto a kernel thread. When you call new Thread(r).start(), the JVM asks the operating system for a thread, which gets its own native stack, kernel bookkeeping and a slot in the kernel run queue. The kernel decides when it runs, preempts it when its time slice ends, and parks it in the kernel when it blocks on a socket or a lock. The JVM is mostly a passenger.
A virtual thread maps many to few. It is a Java object holding a continuation, the captured state of a computation that can be suspended and resumed. To run, it is mounted on a carrier, an ordinary platform thread owned by the JVM's scheduler, a ForkJoinPool in FIFO mode whose parallelism defaults to the number of available processors. When the code blocks on something the JDK knows how to wait for asynchronously, such as a socket read or a ReentrantLock, the JVM unmounts it: its frames are copied to the heap, the carrier goes off to run another virtual thread, and when the I/O or lock is ready the virtual thread is rescheduled, possibly on a different carrier.
Both are instances of Thread; ThreadLocal, interrupts, synchronized and java.util.concurrent all work. The differences sit underneath the API.
What each kind of thread costs
Memory. A platform thread reserves a native stack of a fixed maximum size, set by -Xss or the Thread constructor; on 64-bit Linux the default is typically 1 MB. That is a reservation of address space, and the operating system commits pages only as the stack is touched, so a shallow thread uses far less physical memory than its reservation. The real limits are usually ulimit -u, kernel.threads-max, memory-mapping counts, container PID limits and scheduler overhead, long before RAM.
A virtual thread has no fixed stack. While mounted it runs on the carrier's stack; when it unmounts, the frames it actually uses are copied into heap objects called stack chunks. Its memory is proportional to its real call depth and lives on the garbage-collected heap. A million parked virtual threads with shallow stacks fit in an ordinary heap; with deep framework stacks they will not, and the symptom is GC pressure rather than a thread-creation error.
Creation. Starting a platform thread is a system call plus stack setup, which is why server code pools them. Starting a virtual thread is allocating an object and submitting it to a queue, cheap enough that the intended pattern is one virtual thread per task, never a pool.
Switching. A platform-thread switch is a kernel context switch. A virtual-thread switch is a user-space copy of changed frames to or from the heap: no kernel, but not free, since the copy grows with stack depth. Measure both on your own hardware with the harness below rather than trusting published numbers.
Blocking and scheduling compared
A blocked platform thread costs a kernel thread for as long as it waits. A blocked virtual thread costs a heap object and frees its carrier, provided the JDK can unmount it. Where it cannot, it is pinned or captures its carrier, occupying one of your few carriers.
| Operation | Platform thread | Virtual thread |
|---|---|---|
Socket read or write, HttpClient, JDBC over sockets | Kernel thread blocks | Unmounts; carrier is freed |
ReentrantLock, Semaphore, BlockingQueue | Parks in the kernel | Unmounts |
Thread.sleep, LockSupport.park | Parks in the kernel | Unmounts |
synchronized and Object.wait | Parks in the kernel | JDK 21: pins the carrier. JDK 24 and later (JEP 491): unmounts in general |
| File I/O | Kernel thread blocks | Captures the carrier; the scheduler may temporarily add a carrier to compensate |
| Native code and JNI frames on the stack | Kernel thread blocks | Pinned |
| Class loading or a class initializer | Normal | Pinned, even on JDK 24 and later |
| Long CPU loop | Preempted by kernel time slicing | Not preempted; holds the carrier until it blocks or ends |
The last row is the second big difference: scheduling policy. The kernel time-slices platform threads, so a CPU-heavy thread cannot starve its neighbours for long. The virtual-thread scheduler is cooperative at blocking points: a virtual thread runs until it blocks or finishes. Eight CPU-bound virtual threads on eight cores occupy every carrier while thousands of I/O-bound ones wait. Virtual threads also ignore setPriority and are always daemon threads. Common real-world traps are catalogued in virtual thread pitfalls.
Same code, both models: a harness
The API is deliberately the same, so one program can exercise both models. This harness starts N tasks that each block for a fixed time and reports wall time and peak live threads. Run it with platform and virtual threads at increasing N, and with the stack depth and sleep time adjusted to resemble your service.
// ThreadModels.java - run: java ThreadModels virtual 100000 or java ThreadModels platform 2000
import java.time.Duration;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicInteger;
public class ThreadModels {
static final AtomicInteger live = new AtomicInteger();
static final AtomicInteger peak = new AtomicInteger();
static void task() {
int now = live.incrementAndGet();
peak.accumulateAndGet(now, Math::max);
try {
Thread.sleep(Duration.ofMillis(200)); // stands in for a downstream call
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
} finally {
live.decrementAndGet();
}
}
public static void main(String[] args) {
boolean virtual = args[0].equals("virtual");
int n = Integer.parseInt(args[1]);
long start = System.nanoTime();
try (ExecutorService ex = virtual
? Executors.newVirtualThreadPerTaskExecutor()
: Executors.newFixedThreadPool(200)) { // a typical server pool
for (int i = 0; i < n; i++) ex.submit(ThreadModels::task);
} // close() waits for every task
long ms = (System.nanoTime() - start) / 1_000_000;
System.out.println(args[0] + " n=" + n + " wall_ms=" + ms + " peak_live=" + peak.get());
}
}With a pool of 200, wall time grows linearly with N: 10,000 tasks take about 50 rounds of 200 ms. With virtual threads, peak concurrency equals N and wall time stays near 200 ms until heap becomes the limit. Virtual threads remove the thread limit, not the work.
Sizing with queueing arithmetic: a worked example
Whether virtual threads help is arithmetic. Little's law says the average number of requests in flight equals arrival rate times time in system: L = λ × W. In a thread-per-request server each in-flight request holds a thread, so L is the number of threads you need.
Worked example. A checkout service receives 2,000 requests per second. Each takes 150 ms, of which 10 ms is CPU and 140 ms is waiting on a payment API and a database. Threads needed: 2,000 × 0.15 = 300. CPU needed: 2,000 × 0.010 = 20 core-seconds per second, so 20 cores. A 300-thread platform pool handles this; virtual threads change little.
Now the payment API degrades and its latency rises to 2 seconds. W becomes about 2.01 s and L becomes 4,020. The 300-thread pool saturates and latency explodes for every endpoint sharing it, including ones that never call payments. With virtual threads, 4,020 threads are cheap and the service keeps accepting work, but 4,020 concurrent calls now hit a payment API that is already struggling.
So the rule is: virtual threads help when L is large because W is dominated by waiting, and CPU is not the bottleneck. They do not help when work is CPU-bound, because L is then bounded by cores, and they move rather than remove the bottleneck when a downstream has a fixed capacity. The platform pool used to act as an accidental limit on downstream load; after migration you must put that limit back explicitly.
Migrating a thread-pool service
A safe migration keeps the code and replaces implicit pool limits with explicit ones.
// Before: the pool size is both a thread budget and a hidden limit on downstream load.
ExecutorService workers = Executors.newFixedThreadPool(200);
// After: one virtual thread per task, and the limit made explicit per dependency.
ExecutorService workers = Executors.newVirtualThreadPerTaskExecutor();
Semaphore paymentPermits = new Semaphore(64); // sized from the payment API's capacity
PaymentResult charge(Order o) throws InterruptedException {
if (!paymentPermits.tryAcquire(50, TimeUnit.MILLISECONDS)) {
throw new RejectedExecutionException("payment bulkhead full"); // shed load early
}
try {
return paymentClient.charge(o); // blocking call; the virtual thread unmounts
} finally {
paymentPermits.release();
}
}
// CPU-heavy work stays on a bounded platform pool so it cannot monopolise carriers.
ExecutorService cpuPool = Executors.newFixedThreadPool(Runtime.getRuntime().availableProcessors());In frameworks the switch is usually a setting: Spring Boot 3.2 and later run request handling on virtual threads with spring.threads.virtual.enabled=true, and recent Tomcat and Jetty versions offer virtual-thread executors. Underneath: never pool virtual threads, keep connection pools sized to the database, and audit ThreadLocal caches, which become per-request allocations. For request context that used ThreadLocal, ThreadLocal versus ScopedValue covers the move to ScopedValue, final since JDK 25. A rollout plan with canaries and telemetry is in virtual threads in production.
Observability differences
Tools built for platform threads see virtual threads only partially, and the difference matters during an incident.
- Thread dumps.
jstackandThread.getAllStackTraces()show platform threads, including carriers, but not parked virtual threads. Usejcmd <pid> Thread.dump_to_file -format=json <file>, which includes virtual threads grouped by container, and group the result by stack to find the lock or pool they are waiting on. - Thread counts.
ThreadMXBeancounts platform threads. A dashboard alarm on thread count goes quiet after migration, so replace it with in-flight request counts and permits in use. - JFR.
jdk.VirtualThreadPinnedrecords pinning above a threshold,jdk.VirtualThreadSubmitFailedrecords scheduling failures, and the start and end events can be enabled for sampling. This is the main source of truth for carrier problems.
Failure modes side by side
| Failure | Platform-thread form | Virtual-thread form |
|---|---|---|
| Too much concurrency | Pool saturates; requests queue; RejectedExecutionException or timeouts | No queue in front; downstream pools and APIs overload; heap grows with parked stacks |
| Slow dependency | All shared-pool endpoints stall | Only callers of that dependency stall, if each has its own bulkhead |
| CPU hog | Kernel time slicing keeps others responsive | Carriers monopolised; I/O-bound threads wait despite idle CPUs elsewhere |
| Blocking that cannot unmount | Not applicable | Pinned carriers; on JDK 21 mostly synchronized I/O, on 24 and later native frames and class initialisation |
| Leaked threads | Visible in jstack; capped by the pool | Invisible to jstack; unbounded; found only in JSON dumps or heap |
| Per-thread caches | 200 copies, fine | One copy per request, allocation and GC churn |
Trade-offs and a decision matrix
Neither model is better in general; each is right for a shape of work.
| Choose | When |
|---|---|
| Virtual threads | Request-per-thread servers and clients whose time is mostly waiting on network I/O; fan-out to many services; code you want to keep synchronous and readable instead of rewriting into callbacks or reactive chains |
| Platform threads | CPU-bound computation, where cores are the limit; code that blocks in native libraries or JNI; work that needs priorities, time slicing or a dedicated thread, such as a game loop or audio thread; JDK 21 code with I/O under synchronized that you cannot change |
| Both | Most real services: virtual threads for request handling and I/O, a small bounded platform pool for CPU-heavy steps, and explicit semaphores in front of every dependency |
Compared with reactive code, virtual threads keep ordinary stack traces and exception handling but give up explicit backpressure, which you rebuild with semaphores. Compared with a tuned pool, they remove saturation failures but move overload downstream. For splitting work inside a request, structured concurrency pairs naturally with virtual threads; it remained a preview API through JDK 25, so check its status in your JDK before depending on it.
What to do next
- Measure L for your service with Little's law: arrival rate times average latency, split into CPU time and wait time. If L is in the hundreds and CPU-bound, stop here; a platform pool is fine.
- Run the harness above on your production JDK and hardware with your real stack depth, and record wall time, heap and peak concurrency for both models.
- Run JDK 25 if you can, so that
synchronizedno longer pins; on JDK 21, find I/O undersynchronizedwith thejdk.VirtualThreadPinnedJFR event. - Before switching, inventory every limit the old pool enforced implicitly, and add a semaphore or bulkhead per dependency sized from that dependency's capacity.
- Move CPU-heavy steps to a bounded platform pool, and remove any pooling of virtual threads.
- Replace thread-count dashboards with in-flight requests, permits in use, JFR pinning events and JSON thread dumps grouped by stack.
- Canary the switch on one service, compare tail latency and downstream error rates against the platform-thread baseline, and keep the old executor behind a flag until the numbers hold under a dependency slowdown.