Project Loom is the OpenJDK effort that made concurrency cheap enough to write in the plain blocking style again. It is not one feature but a family: virtual threads, the rework of synchronized so it no longer pins, scoped values, and structured concurrency. Some of those are final and some are still previews, and most confusion about Loom comes from mixing up which is which.
This page takes the project-level view. It explains from first principles why threads were the bottleneck, shows the mechanism in enough detail to predict behaviour, walks a real migration of a service from a fixed thread pool to one virtual thread per task, and lists what breaks. The per-feature deep dives are linked where they go further.
What shipped, and when
Loom's pieces arrived over several releases. Check the JDK you run against this table before copying code from a blog post, because preview APIs have changed shape between releases.
| Feature | JEP | Status |
|---|---|---|
| Virtual threads | 425 (JDK 19), 436 (JDK 20), 444 (JDK 21) | Previews, then final in JDK 21 |
| Synchronize virtual threads without pinning | 491 | JDK 24; synchronized and Object.wait() no longer pin |
| Scoped values | 506 (after several previews) | Final in JDK 25 |
| Structured concurrency | 505 (JDK 25), 525 (JDK 26) | Still preview; fifth and sixth previews, API revised between them |
The practical reading: on JDK 21 you have virtual threads but synchronized blocks can pin; on JDK 25, the current long-term-support release, pinning is mostly gone and scoped values are final; structured concurrency needs --enable-preview on every release so far and its API should be treated as unstable.
Why threads were the bottleneck
Little's law says that the average number of requests in flight equals throughput multiplied by latency. A service that handles 2,000 requests per second where each request spends 200 ms mostly waiting on a database and two HTTP calls has 400 requests in flight at any moment. In the thread-per-request style each one occupies a thread for its whole life, waiting included.
A platform thread is a thin wrapper over an operating-system thread. Each reserves a stack (commonly 1 MB of address space on 64-bit Linux, committed lazily), costs a kernel data structure, and is scheduled by the kernel. Thousands are fine; hundreds of thousands are not. So services capped their pools, typically at a few hundred. With 200 threads and 200 ms of latency the ceiling is 200 / 0.2 = 1,000 requests per second, no matter how idle the CPUs are. When a downstream slows to 2 s the ceiling falls to 100, the queue grows, and the service falls over while the CPU sits at 10 percent.
The industry's answer was asynchronous code: callbacks, CompletableFuture chains, reactive streams. They decouple waiting from threads, but they split one logical request across many stack traces, break try/catch and debuggers, and make thread-local context unreliable. Loom's bet was the reverse: keep the blocking code and make the thread cheap.
The mechanism: continuations on carrier threads
A virtual thread is a java.lang.Thread whose execution state lives in a heap object called a continuation. The JDK runs it by mounting it on a carrier, an ordinary platform thread owned by a scheduler. When the code reaches a blocking operation that the JDK has adapted, such as a socket read, a ReentrantLock wait or a BlockingQueue.take(), the virtual thread parks: its stack frames are copied into stack-chunk objects on the heap, it unmounts, and the carrier picks up other work. When the event arrives (the poller sees the socket is readable, or the lock is released) the virtual thread is resubmitted and resumes, possibly on a different carrier.
The default scheduler is a work-stealing ForkJoinPool in FIFO mode. Its size is controlled by jdk.virtualThreadScheduler.parallelism, which defaults to the number of available processors, and jdk.virtualThreadScheduler.maxPoolSize, which defaults to the larger of parallelism and 256. The extra headroom is not for normal operation: the scheduler adds carriers temporarily only to compensate when a carrier is captured by an operation that cannot unmount, such as some file-system calls.
Two consequences follow. First, a virtual thread is cheap to create and to block, so the right unit is one virtual thread per task, never a pool of them. Second, virtual threads do not make CPU work faster; with parallelism equal to the core count, eight compute-bound virtual threads on an eight-core machine are as fast as eight platform threads and no faster. The win is entirely in overlapping waits. The virtual threads architecture page walks each component in more detail.
Pinning, before and after JEP 491
A virtual thread is pinned when it blocks but cannot unmount, so it holds its carrier hostage. On JDK 21 the common cause was blocking inside a synchronized block or method, because the monitor implementation tied ownership to the platform thread. A service with eight carriers where eight requests block inside synchronized on a slow call has no carriers left, and every other virtual thread stalls. That is why JDK 21 guidance was to replace hot synchronized sections with ReentrantLock.
JEP 491, delivered in JDK 24, reimplemented monitors so a virtual thread can acquire, hold and release a monitor independently of its carrier. Blocking to enter a monitor, blocking while holding one, and Object.wait() now unmount. The remaining pinning cases involve native frames on the stack: resolving symbolic references and loading classes, blocking inside a class initializer, and waiting for another thread to finish initializing a class. Calls into native code through JNI or the foreign function API also pin while they run.
Diagnostics changed with it. The jdk.tracePinnedThreads system property was removed and setting it has no effect. The JFR event jdk.VirtualThreadPinned remains and now reports the pinning reason and the carrier. If you are still on JDK 21, the property works there and the old advice to swap synchronized for explicit locks around blocking calls still applies.
Worked example: migrating a pool-per-request service
Take an order-page endpoint that loads a customer, their recent orders and a recommendations call. The pre-Loom version fans out onto a fixed pool and joins the futures:
// Before: fixed pool, sized by guesswork, shared by every request
private final ExecutorService pool = Executors.newFixedThreadPool(200);
OrderPage load(long customerId) throws Exception {
Future<Customer> c = pool.submit(() -> customers.find(customerId));
Future<List<Order>> o = pool.submit(() -> orders.recent(customerId, 20));
Future<List<Item>> r = pool.submit(() -> recs.forCustomer(customerId));
return new OrderPage(c.get(2, SECONDS), o.get(2, SECONDS), r.get(2, SECONDS));
}The migration has three steps. Switch the server to run each request on a virtual thread (Tomcat, Jetty, Helidon and Spring Boot all have a switch; in Spring Boot 3.2 and later it is spring.threads.virtual.enabled=true). Replace the shared pool with a per-task executor. Then move the concurrency limit from the thread count, where it used to live by accident, to the resource that actually needs protecting:
// After: one virtual thread per subtask; the limit protects the dependency, not the JVM
private static final Semaphore RECS_PERMITS = new Semaphore(50);
OrderPage load(long customerId) throws Exception {
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
Future<Customer> c = exec.submit(() -> customers.find(customerId));
Future<List<Order>> o = exec.submit(() -> orders.recent(customerId, 20));
Future<List<Item>> r = exec.submit(() -> {
RECS_PERMITS.acquire();
try { return recs.forCustomer(customerId); }
finally { RECS_PERMITS.release(); }
});
return new OrderPage(c.get(2, SECONDS), o.get(2, SECONDS), r.get(2, SECONDS));
} // close() waits for all subtasks
}Run the arithmetic again. At 2,000 requests per second and 200 ms, 400 requests and up to 1,200 subtasks are in flight; that is a few megabytes of heap for stacks rather than a hard ceiling. When the recommendations service slows to 2 s, in-flight subtasks rise to several thousand, which virtual threads absorb, and the semaphore caps the pressure on the slow dependency at 50 concurrent calls. The database connection pool, not the thread pool, now bounds database concurrency, so size it deliberately and give getConnection() a timeout.
What breaks when threads become cheap
- Thread-local caches. Code that caches an expensive object per thread, such as a 64 KB buffer or a formatter, assumed a few hundred threads. With a million virtual threads it allocates a million copies and each is garbage after one request. Replace caches with pooled or immutable objects, and see the ThreadLocal guide for the trade-offs.
- Pooling virtual threads. Wrapping them in a fixed pool reintroduces the old ceiling. Limit concurrency with a semaphore around the resource instead.
- Unbounded fan-out. The thread pool used to be an implicit back-pressure mechanism. Without it, a traffic spike reaches the database in full. Every external dependency needs an explicit limit and timeout.
- CPU-heavy work. Long computations on virtual threads occupy carriers and starve I/O-bound threads. Put heavy compute on a bounded platform-thread pool.
- Native and legacy blocking. JNI calls, some drivers that block in native code, and class initialization still pin. Profile with JFR before assuming the problem is solved.
- Thread identity assumptions. Metrics labelled by thread name, per-thread MDC logging contexts that are never cleared, and code that expects
Thread.currentThread()to be long-lived all need review.
Scoped values and structured concurrency
Cheap threads make two older patterns awkward. Inheritable thread locals copy context into every child, which is expensive with many children and mutable in ways that surprise. Scoped values, final in JDK 25, bind an immutable value for the duration of a call and make it visible to callees and to child threads forked inside structured scopes:
static final ScopedValue<RequestContext> CTX = ScopedValue.newInstance();
void handle(Request req) {
ScopedValue.where(CTX, RequestContext.from(req))
.run(() -> router.dispatch(req)); // CTX.get() valid anywhere below
}Structured concurrency makes a group of subtasks a single unit: they are forked inside a scope, the scope joins them, and a failure cancels the siblings. It is still a preview. The following compiles on JDK 25 with --enable-preview; JDK 26 revised the joiner methods, so check the javadoc for your release:
// JDK 25, --enable-preview (JEP 505). The default joiner fails fast.
try (var scope = StructuredTaskScope.open()) {
Subtask<Customer> c = scope.fork(() -> customers.find(id));
Subtask<List<Order>> o = scope.fork(() -> orders.recent(id, 20));
scope.join(); // throws if any subtask failed; others are cancelled
return new OrderPage(c.get(), o.get(), List.of());
}Compared with the executor version, a failure in one subtask interrupts the others immediately, thread dumps show the parent-child relationship, and no subtask can outlive the method. Details are on the structured concurrency and scoped values pages.
Operating Loom in production
Observe before and after the migration with the same load test. The useful signals are: JFR events jdk.VirtualThreadPinned (pins longer than a threshold) and jdk.VirtualThreadSubmitFailed; heap growth from stack chunks when millions of threads park; carrier utilisation; and latency percentiles of each dependency. A classic jstack shows platform threads but not the parked virtual ones, so use the JSON thread dump introduced with JEP 444:
jcmd <pid> Thread.dump_to_file -format=json /tmp/threads.json
java -XX:StartFlightRecording=duration=120s,filename=loom.jfr -jar app.jar
jfr print --events jdk.VirtualThreadPinned loom.jfrLeave the scheduler properties at their defaults unless measurements say otherwise. Raising parallelism above the core count rarely helps; if carriers are exhausted the cause is pinning or CPU-bound work, and more carriers only hide it.
Trade-offs against reactive code
Reactive frameworks still have strengths: explicit back-pressure between stages, streaming operators, and mature integration where the whole stack is non-blocking. Loom's strengths are readable sequential code, real stack traces, normal exception handling and debugger stepping, and compatibility with the vast body of blocking libraries such as JDBC. For request-response services dominated by waiting on I/O, virtual threads usually give comparable throughput with much simpler code. For streaming pipelines with fine-grained back-pressure, reactive libraries remain a reasonable choice, and the two can coexist.
What to do next
- Confirm your JDK: plan on 21 at minimum and prefer 25 or later so
synchronizeddoes not pin. - Write down each endpoint's throughput and latency and compute in-flight requests with Little's law; that tells you whether threads were your ceiling at all.
- Enable virtual threads in your server framework behind a flag and replace shared task pools with
newVirtualThreadPerTaskExecutor(). - Add a semaphore and timeout in front of every external dependency, and size connection pools deliberately.
- Audit thread-local caches, MDC handling and CPU-heavy code paths; move heavy compute to a bounded platform pool.
- Load-test with JFR recording on, check
jdk.VirtualThreadPinned, and compare heap, latency percentiles and error rates with the old build before rolling out.