volatile is the lightest synchronization tool in Java and the most often misused. It makes reads and writes of one field visible across threads and orders them with respect to the surrounding code, so a value written by one thread can safely publish everything that thread did before the write. It does not make compound operations atomic, does not protect the contents of an object or array the field points to, and does not block. Those three limits explain nearly every volatile bug.
This article builds the guarantee from the Java Memory Model rather than from the old "reads go to main memory" folklore, which is wrong about how caches work and misleading about what is actually reordered. Then it shows what HotSpot emits on x86 and ARM, works through the classic patterns, introduces the finer-grained VarHandle access modes, and ends with a test you can run to see the difference.
What actually reorders memory operations
Three layers can make one thread's writes appear late or out of order to another. The javac compiler does very little reordering, but the JIT compilers (C2 and Graal) do a lot: they keep values in registers, hoist loop-invariant reads out of loops, and reorder independent memory operations. The CPU adds its own reordering: every modern core has a store buffer, so a store can sit in the buffer while a later load from a different address completes, which looks to other cores like the load happened first. ARM and POWER also reorder loads with loads and stores with stores; x86 does not.
Caches are not the problem people think. Hardware cache coherence (MESI and its variants) keeps cache lines consistent between cores; once a store leaves the store buffer, other cores see it. What breaks programs is the order in which operations become visible and the compiler's freedom to never perform a read at all. volatile constrains exactly those two things.
What actually reorders memory operations
Three layers can make one thread's writes appear late or out of order to another. The javac compiler does very little reordering, but the JIT compilers (C2 and Graal) do a lot: they keep values in registers, hoist loop-invariant reads out of loops, and reorder independent memory operations. The CPU adds its own reordering: every modern core has a store buffer, so a store can sit in the buffer while a later load from a different address completes, which looks to other cores like the load happened first. ARM and POWER also reorder loads with loads and stores with stores; x86 does not.
Caches are not the problem people think. Hardware cache coherence (MESI and its variants) keeps cache lines consistent between cores; once a store leaves the store buffer, other cores see it. What breaks programs is the order in which operations become visible and the compiler's freedom to never perform a read at all. volatile constrains exactly those two things.
The guarantee in Java Memory Model terms
The Java Memory Model (JLS chapter 17) defines correctness through the happens-before relation. Within one thread, each action happens-before the next in program order. Across threads, edges come from synchronization: unlocking a monitor happens-before every later lock of it, starting a thread happens-before its first action, and, the rule that matters here, a write to a volatile field happens-before every subsequent read of that field that sees the value. Happens-before is transitive, so everything thread A did before its volatile write is visible to thread B after B's volatile read observes that write.
Two further properties are worth knowing. All volatile accesses take part in a single total synchronization order consistent with each thread's program order, so volatile reads and writes are sequentially consistent among themselves. And a program with no data races, where every conflicting pair of plain accesses is ordered by happens-before, behaves as if it were sequentially consistent. That is the practical goal: use volatile, locks or atomics to remove the races, then reason as if operations ran one at a time. The Java Memory Model article covers the full rule set, including final field semantics.
Worked example: the stop flag
The smallest useful example is a stop flag:
class Worker implements Runnable {
private boolean running = true; // BROKEN: plain field
public void run() { while (running) { step(); } }
public void stop() { running = false; }
}If step() is inlined and never writes running, C2 is allowed to read the field once, keep it in a register and compile the loop as if (running) while (true) step();. The write from stop() reaches memory, but the worker never reads memory again. In practice this shows up only after the method is JIT-compiled, so it passes in a debugger and hangs in production. Declaring private volatile boolean running forbids the hoisting: every iteration performs a real load. A single writer setting a flag that readers poll is the textbook volatile use, because the new value does not depend on the old one.
Worked example: publishing a configuration snapshot
The more valuable use is safe publication of an object graph. Suppose a service reloads its configuration every minute and request threads read it constantly:
final class Config {
final Duration timeout;
final List<String> hosts;
Config(Duration timeout, List<String> hosts) {
this.timeout = timeout;
this.hosts = List.copyOf(hosts);
}
}
class ConfigHolder {
private volatile Config current = load();
Config get() { return current; } // one volatile load per request
void reload() { // called by one scheduler thread
Config next = load(); // build fully, privately
current = next; // volatile store publishes it
}
}The diagram shows the edge. The constructor's writes happen-before the volatile store in the scheduler thread; the store happens-before any request thread's load that sees the new reference; so the request thread sees fully built timeout and hosts values. Readers never lock, and a reload is a single reference swap. Two conditions keep it correct: the object must not be modified after publication (make it immutable, as here), and each request should read get() once into a local and use that local, so it does not mix two configurations within one request.
Visibility is not atomicity
The trap is read-modify-write. count++ on a volatile int is a volatile read, an add and a volatile write. Two threads can both read 41 and both write 42. Each access is visible and ordered, and an increment is still lost. Run two threads doing a million increments each on a volatile counter and the total typically falls well short of two million. The same applies to check-then-act sequences such as if (cache == null) cache = build();, which can build twice.
Use AtomicInteger.incrementAndGet() or compareAndSet for one variable that must be updated atomically, LongAdder for a hot counter that many threads increment and few read, and a lock when the invariant spans several fields. The atomic classes article explains the compare-and-set loop these are built on.
Double-checked locking and the holder idiom
Double-checked locking is correct only with volatile:
class Registry {
private static volatile Registry instance;
static Registry get() {
Registry r = instance; // one volatile read on the fast path
if (r == null) {
synchronized (Registry.class) {
r = instance;
if (r == null) {
instance = r = new Registry();
}
}
}
return r;
}
}Without volatile, the store of the reference can become visible before the constructor's stores, and a thread on the unlocked fast path can return a partially initialised object. The local variable r avoids a second volatile read. For static singletons, the holder idiom is simpler and needs no volatile, because class initialization is already guarded by the JVM: a nested static class Holder whose static final field creates the instance. Reserve double-checked locking for lazily created instance fields, and see synchronized for the locking half.
What HotSpot emits, and what it costs
On x86-64, which already keeps loads and stores in order except for store-to-later-load, HotSpot compiles a volatile load to an ordinary move and a volatile store to a move followed by a StoreLoad barrier, implemented as a lock-prefixed instruction on a stack location. The barrier drains the store buffer, so volatile writes cost noticeably more than plain ones and volatile reads cost almost nothing beyond the lost optimisations. On AArch64 HotSpot uses the load-acquire ldar and store-release stlr instructions, whose combination also provides the sequentially consistent ordering volatile needs. In both cases the larger cost is often the JIT optimisations a volatile access prevents: no register caching, no hoisting, no merging of repeated reads.
Two related details. Under JLS 17.7, writes to plain long and double fields may be split into two 32-bit halves on some JVMs; volatile ones are always atomic. And a volatile field that is written heavily by several threads suffers from cache-line contention; if neighbouring hot fields share its line you get false sharing, which the JDK avoids internally with the @Contended annotation (usable by application code only with -XX:-RestrictContended).
VarHandle access modes
Since Java 9, VarHandle exposes weaker and cheaper orderings for a field, array element or buffer location. From weakest to strongest: plain access; opaque, which is never hoisted or torn but gives no ordering with other variables; acquire/release, which gives exactly the publication edge from the diagram without the StoreLoad barrier; and volatile, identical to the keyword.
class Mailbox {
private Message slot; // plain field, accessed through the handle
private static final VarHandle SLOT;
static {
try {
SLOT = MethodHandles.lookup().findVarHandle(Mailbox.class, "slot", Message.class);
} catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
}
void put(Message m) { SLOT.setRelease(this, m); } // publish without a full fence
Message peek() { return (Message) SLOT.getAcquire(this); }
boolean claim(Message expected) { return SLOT.compareAndSet(this, expected, null); }
}Release/acquire is enough for single-producer handoffs and is what many queue implementations use. It is not enough for Dekker-style protocols where each thread writes one variable and then reads another; those need full volatile semantics. VarHandles also solve the volatile-array problem: a volatile array field makes the reference volatile, not its elements, so use MethodHandles.arrayElementVarHandle or AtomicIntegerArray for element-level ordering. The method handles article covers lookup and access rules.
Testing ordering with jcstress
Memory-ordering bugs rarely reproduce in unit tests. The OpenJDK jcstress harness runs small actors concurrently millions of times and tallies outcomes. This test checks the store-buffer case:
@JCStressTest
@Outcome(id = {"0, 1", "1, 0", "1, 1"}, expect = Expect.ACCEPTABLE, desc = "sequentially consistent")
@Outcome(id = "0, 0", expect = Expect.FORBIDDEN, desc = "store-load reordering")
@State
public class DekkerVolatile {
volatile int x, y;
@Actor public void actor1(II_Result r) { x = 1; r.r1 = y; }
@Actor public void actor2(II_Result r) { y = 1; r.r2 = x; }
}With volatile fields the 0, 0 outcome never appears. Remove the keyword and, on x86 hardware, 0, 0 shows up quickly, because each core's store is still in its store buffer when it loads the other variable. Keep tests like this next to any hand-written lock-free code.
Choosing the right tool, and failure modes
| Need | Use | Why |
|---|---|---|
| Stop flag, status, one writer | volatile | visibility, no blocking |
| Publish an immutable snapshot | volatile reference | happens-before edge |
| Single-producer handoff, hot path | VarHandle release/acquire | no full fence |
| Atomic update of one variable | AtomicInteger, AtomicReference | CAS |
| Hot counter, rare reads | LongAdder | spreads contention |
| Invariant across fields | synchronized or a lock | mutual exclusion |
| Shared map | ConcurrentHashMap | built-in safe publication |
Failure modes to look for in review: a volatile counter incremented from several threads; a volatile reference to a mutable collection that is modified after publication; a volatile array whose elements are expected to be ordered; reading a volatile field twice and assuming both reads match; and volatile added "for safety" to fields guarded by a lock, which costs performance and hides the real design.
What to do next
- List every field shared between threads and write down what guards it: final, volatile, a lock or an atomic.
- Convert flags and published snapshots to volatile; make the snapshot classes immutable.
- Replace any
++or check-then-act on a volatile with an atomic class or a lock. - Replace static double-checked locking with the holder idiom.
- Use VarHandle release/acquire only where profiling shows the fence matters, and document why.
- Add a jcstress test for every hand-written lock-free structure.
- Read the JIT output with
-XX:+PrintAssembly(needs the hsdis plugin) when you need to confirm barriers.