A thread-local variable gives each thread its own copy of a value behind one shared name. Code deep in a call stack can read the current request id, transaction, security principal or reusable buffer without every method in between passing it along. That convenience is why logging, transaction management and tracing libraries all lean on it, and why its failure modes show up in production rather than in tests.
This article is about how thread-local storage is built and how to carry the context it holds across the places where threads change. It opens Java's ThreadLocalMap and walks through its hashing and cleanup, goes down to the operating system and CPU register that make native thread-locals cheap, compares the designs other runtimes chose, and ends with the capture-and-restore pattern that executors and async code need. The pool leak and InheritableThreadLocal traps are covered in the Java ThreadLocal article, and the comparison with ScopedValue in ThreadLocal vs ScopedValue.
Three ways to build per-thread storage
There are only three basic designs. A global map keyed by thread identity is the obvious one, but every access needs synchronisation on a shared structure and entries for dead threads must be removed by someone. A per-thread map keyed by variable inverts it: each thread owns a private table, so reads and writes need no locks because only the owning thread touches its table. The third design pushes the work into the compiler, linker and operating system: each thread gets a block of memory laid out at load time, and a variable is a fixed offset from a per-thread base address.
Java uses the second design for ThreadLocal, and the JVM itself relies on the third for its own bookkeeping. C, C++ and Rust expose the third directly. Knowing which design you are using explains its costs: per-thread maps cost a hash lookup and some memory per thread, while native thread-locals cost an add to a register.
Inside ThreadLocalMap
Every java.lang.Thread has a threadLocals field and an inheritableThreadLocals field, each pointing to a ThreadLocalMap that is created lazily on first use. The map is a custom open-addressing hash table, not a HashMap: an array of Entry objects with an initial capacity of 16, always a power of two, resolving collisions by linear probing to the next slot.
Each Entry extends WeakReference<ThreadLocal<?>>. The key, the ThreadLocal object, is held weakly, while the value is an ordinary strong field. The intent is that when application code drops its last reference to a ThreadLocal, the garbage collector can clear the key, leaving a stale entry whose key reads as null. The value is still strongly reachable from the thread until something cleans the slot.
Calling get() fetches the current thread, reads its map, computes the slot and checks whether the entry there refers to this key. On a hit it returns the value. On a miss it probes forward, and if nothing is found it calls initialValue() (or the supplier given to ThreadLocal.withInitial) and stores the result, so a get can allocate and insert.
The golden-ratio hash
A ThreadLocal does not use hashCode() for its slot. Each instance takes a threadLocalHashCode from a global AtomicInteger that advances by 0x61c88647 per new instance. That constant is about 2 to the 32 divided by the golden ratio, and the JDK comment explains the choice: it turns sequential ids into near-optimally spread multiplicative hash values for power-of-two tables. Consecutively created thread-locals, the common case, land far apart instead of colliding.
public class SlotDemo {
public static void main(String[] args) {
final int HASH_INCREMENT = 0x61c88647;
int h = 0;
for (int i = 0; i < 8; i++) {
System.out.printf("local %d -> slot %2d of 16%n", i, h & 15);
h += HASH_INCREMENT;
}
}
}
// slots: 0, 7, 14, 5, 12, 3, 10, 1 -- no collisions among the first eightThe mask h & (len - 1) replaces a modulo because the length is a power of two. With linear probing, good spread matters more than with chaining, because one collision can lengthen the probe sequence for every key whose home slot lies after it.
Lazy cleanup: expunging, threshold and resize
No background thread scans the map. Cleanup piggybacks on normal operations. When a probe meets an entry whose key has been cleared, expungeStaleEntry nulls its value, removes it and re-inserts the following entries in the same run so probe chains stay intact. After a set inserts a new entry, cleanSomeSlots checks a logarithmic number of further slots, halving its counter each step, which bounds the cost of an insert while still finding garbage eventually. remove() clears the entry and expunges immediately.
The resize policy has two stages. The threshold is two thirds of the table length. When an insert reaches it and the light scan found nothing to clean, rehash() first expunges every stale entry in the table, and only if the size is still at least three quarters of the threshold, which is half the table, does resize() double the array. The result is a table that stays sparse and short-probed for the handful of thread-locals most threads use, at the price of leaving stale values in place on threads that never touch their map again.
That last point is the architectural limit of weak keys. A pooled thread that is idle, or that only ever uses a few fixed thread-locals, may never probe the slot holding a large stale value, so the value lives as long as the thread. Weak keys make leaks less likely; only remove() makes them impossible.
Underneath: native thread-local storage
Native code gets per-thread data from the operating system and the CPU. On x86-64 Linux the thread pointer lives in the base of the fs segment, and on AArch64 in the TPIDR_EL0 system register. For an executable, the linker can place each thread-local variable at a fixed offset from that pointer, so an access is one register-relative load. This is the local-exec model of ELF thread-local storage. Shared libraries loaded at run time may need the general-dynamic model, which calls __tls_get_addr to find the module's block for the current thread, so thread-locals in a dynamically loaded library cost more than those in the main program.
// C++: one counter per thread, no locks needed
thread_local long requests_handled = 0;
void handle() { ++requests_handled; }
// Rust: thread_local! with interior mutability
use std::cell::Cell;
thread_local! { static HANDLED: Cell<u64> = Cell::new(0); }
fn handle() { HANDLED.with(|h| h.set(h.get() + 1)); }The JVM uses the same idea for itself. HotSpot on x86-64 keeps a pointer to the current thread's internal structure in a dedicated register, r15, while running compiled Java code, which is why Thread.currentThread() is close to free and why the first step of ThreadLocal.get() is cheap. The hash lookup that follows is the real cost, and it is still only a few memory accesses.
How other runtimes answer the same question
| Runtime | Per-thread mechanism | Context across async hops |
|---|---|---|
| C and C++ | _Thread_local / thread_local on native TLS | None built in; pass context explicitly |
| Rust | thread_local! with LocalKey::with | Async runtimes provide task-local storage |
| Go | None by design; goroutines have no exposed identity | context.Context passed as the first argument |
| Python | threading.local() | contextvars; asyncio runs each task in a copy of the context captured at creation |
| .NET | ThreadLocal<T>, [ThreadStatic] | AsyncLocal<T> flows with the execution context |
| Java | ThreadLocal, InheritableThreadLocal | ScopedValue (final in JDK 25), or explicit capture |
The pattern is clear: every runtime that grew serious async support ended up with two mechanisms, one tied to the operating-system thread and one tied to the logical flow of work. Go skipped the first entirely to force the second. When you choose where to store context, choose based on which of those two things it really belongs to.
Carrying context across executors
A thread-local is invisible to any task that runs on another thread, which is every task submitted to a pool or a completion stage. The fix is capture and restore: read the values when the task is created, install them when it runs and put back whatever was there when it finishes. Logging diagnostic contexts and tracing libraries all implement this; OpenTelemetry's Java context, for example, is stored in a thread-local by default and offers wrappers for executors.
public final class ContextAwareExecutor implements Executor {
private static final ThreadLocal<String> REQUEST_ID = new ThreadLocal<>();
private final Executor delegate;
public ContextAwareExecutor(Executor delegate) { this.delegate = delegate; }
@Override public void execute(Runnable task) {
final String captured = REQUEST_ID.get(); // on the submitting thread
delegate.execute(() -> {
final String previous = REQUEST_ID.get(); // on the worker thread
REQUEST_ID.set(captured);
try {
task.run();
} finally {
if (previous == null) REQUEST_ID.remove(); else REQUEST_ID.set(previous);
}
});
}
}Restoring the previous value instead of always removing matters when a task runs inline on the caller's thread, for example with a caller-runs rejection policy, because removing would wipe the caller's own context. Pools themselves are covered in the thread pool article.
Virtual threads change the arithmetic
Virtual threads support thread-locals, and each virtual thread gets its own map. The mechanism is unchanged; the scale is not. A server that used two hundred platform threads, each caching a 64 KB buffer in a thread-local, held about 12.5 MB. The same code on a million virtual threads, one per request, would allocate the buffer per request and could hold about 61 GB if they were all alive at once, and the cache never hits because each virtual thread is used once. Thread-locals used as per-thread caches are the pattern to remove when adopting virtual threads; thread-locals used as per-request context still work, though ScopedValue is a better fit.
To find them, run with -Djdk.traceVirtualThreadLocals, which prints a stack trace when a virtual thread sets a thread-local value. The virtual threads article covers the scheduling side.
Failure modes and trade-offs
- Leaked context between requests on a pooled thread: the next task reads the previous task's user or tenant. Always clear in a
finallyblock or use a scoped wrapper. - Class-loader retention in application servers: a value whose class came from an undeployed application keeps that loader alive through a live pooled thread.
- Lost context after an async hop: the log line or trace span is missing its id. Use capture-and-restore wrappers on every executor.
- Hidden coupling: a method's behaviour depends on state its signature does not show, which makes tests order-dependent. Reset thread-locals in test setup.
- Memory per thread: harmless with tens of threads, serious with millions.
The trade-off is explicitness against convenience. Passing context as parameters is always correct and always visible, and it is what Go requires. Thread-locals save plumbing in deep frameworks and are the right tool for genuinely per-thread resources such as non-thread-safe formatters or random generators. Scoped, immutable bindings sit between the two. Visibility of the values themselves across threads follows the rules in the memory model article.
What to do next
- Inventory every
ThreadLocalin your code and dependencies, and label each as per-thread resource or per-request context. - Wrap every set in try-finally with
remove(), or replace it with a helper that restores the previous value. - Wrap executors and completion stages so request context is captured on submit and restored on run.
- Before moving to virtual threads, run with
-Djdk.traceVirtualThreadLocalsand remove per-thread caches. - Move immutable per-request context to
ScopedValueon JDK 25 or later. - Add a test that runs two requests on a single-thread pool and asserts the second cannot see the first's context.