Two threads that never share anything need no synchronization at all. The moment they share a counter, a cache, a queue or a flag, you need two guarantees that hardware and compilers do not give you by default: that one thread's half-finished update is never seen or overwritten by another, and that when one thread finishes something, the other actually sees it. Locks, condition variables, semaphores, latches, atomics and channels are all different ways of buying those two guarantees, and each fits a different shape of problem.
This article is the map rather than a tour of one primitive. It explains the two jobs synchronization does, the single rule (happens-before) that every primitive is built on, how to pick a primitive from the shape of your coordination problem, and then builds a small but realistic Java service that uses four different primitives, each for the job it is best at. It ends with failure modes, how to test concurrent code and a checklist.
The map
Two jobs: exclusion and visibility
Synchronization does two separate jobs, and most bugs come from solving one and forgetting the other.
Exclusion makes a group of operations atomic with respect to other threads. The classic example is count = count + 1: it is a load, an add and a store, and two threads can both load 41 and both store 42. A transfer between two accounts, a check-then-insert into a map, and an update of two fields that must agree are the same problem in larger form. Exclusion says: while I am in this region, nobody else touching the same data is.
Ordering and visibility make one thread's writes visible to another, in an order it can rely on. Modern CPUs keep writes in store buffers and caches, and compilers keep values in registers and reorder independent instructions. Without synchronization, thread B can spin forever on a ready flag that thread A set long ago, or see ready == true but a stale value in the data the flag was guarding. No interleaving explains that result; it comes from reordering, which is why race bugs are so hard to reason about by staring at source code.
A third job, waiting, sits on top of these: a thread needs to sleep until some condition holds (the queue is non-empty, five workers have finished, a permit is free) without burning a CPU core and without missing the moment the condition becomes true.
Happens-before: the one rule underneath
Every mainstream memory model (Java, C++11 and later, Go, Rust through C++'s) is defined in terms of one relation: happens-before. If write W happens-before read R, R is guaranteed to see W or something later. If two accesses to the same location are not ordered by happens-before and at least one is a write, you have a data race. In C and C++ a data race is undefined behaviour; in Java the program stays memory-safe, but the racing read may see stale or partially ordered values.
Within one thread, program order gives you happens-before for free. Between threads you only get it from synchronization: a mutex unlock happens-before the next lock of the same mutex; a write to a Java volatile field happens-before every later read that sees it; Thread.start happens-before the started thread's first action; completing a future happens-before a get that returns its value; putting an item in a BlockingQueue happens-before taking it out. The relation is transitive, which is what makes safe publication work: build an object fully, then publish its reference through any of these edges, and the reader sees the fully built object.
This gives a test you can apply to any piece of shared state: for every pair of accesses from different threads where one is a write, name the happens-before edge that orders them. If you cannot name it, you have a race, however unlikely it looks.
The toolbox, sorted by question
Primitives are best sorted by the question they answer, not by how they are implemented.
| Question | Primitive | Typical Java form | Watch out for |
|---|---|---|---|
| Only one thread may run this region | Mutex | synchronized, ReentrantLock | Holding it across I/O; lock ordering |
| Many readers, rare writers | Read-write lock | ReentrantReadWriteLock, StampedLock | Writer starvation; often slower than a mutex for short sections |
| Sleep until a predicate is true | Condition variable | wait/notifyAll, Condition | Must re-check the predicate in a loop |
| At most N at once | Counting semaphore | Semaphore | Leaked permits when release is skipped |
| Wait for one event, once | Latch, future, once-init | CountDownLatch, CompletableFuture | Waiting forever if the event never fires |
| N threads meet, then all continue | Barrier | CyclicBarrier, Phaser | One dead party blocks everyone |
| Update one variable without a lock | Atomic, compare-and-swap | AtomicLong, LongAdder | Compound invariants across two variables |
| Pass data, not share it | Queue, channel | ArrayBlockingQueue | Unbounded queues hide overload |
| Swap a whole consistent view | Immutable snapshot + volatile ref | volatile, AtomicReference | Mutating the snapshot after publishing it |
Choosing a primitive
A short decision procedure covers most code:
- Can you avoid sharing? Confine state to one thread (an actor, an event loop, a per-request object) or make it immutable. Immutable data needs only safe publication, never locking.
- Is the shared thing a single value? Use an atomic. Counters with heavy write contention do better with
LongAdder, which spreads updates over cells and sums them on read. - Is it a consistent multi-field view that changes rarely? Build a new immutable object and swap the reference. Readers pay one volatile read and never block.
- Is it producer-consumer? Use a bounded queue or channel and let it carry both the data and the happens-before edge.
- Do several fields change together often? Use a mutex, keep the critical section short, and never call unknown code while holding it.
- Do threads need to wait for a state, a count or an event? Use the matching higher-level tool (latch, semaphore, future, barrier) before writing a condition variable by hand.
The order matters. Each step down the list adds a way to deadlock, lose a wakeup or contend. Lock-free algorithms built from many compare-and-swap steps are deliberately missing from the list: they are worth it inside libraries, rarely in application code.
Worked example: a price service with four primitives
Consider a pricing service. It answers price(sku) from a cache, loads missing prices from a slow upstream API, applies a currency table that an operator can reload at any time, and must not answer before its first currency table is loaded. Four requirements, four primitives:
- The currency table is a consistent multi-field view that changes rarely: an immutable snapshot behind a volatile reference.
- When 300 requests for the same uncached SKU arrive together, the upstream must be called once, not 300 times: single-flight loading with
computeIfAbsentand aCompletableFuture. - The upstream tolerates at most 8 concurrent calls: a semaphore as a bulkhead.
- Requests must wait until the first table arrives: a latch.
import java.math.BigDecimal;
import java.util.Map;
import java.util.concurrent.*;
public final class PriceService {
// Immutable view: built fully, then published by one volatile write.
record Rates(Map<String, BigDecimal> perCurrency, long version) {
Rates { perCurrency = Map.copyOf(perCurrency); }
}
private volatile Rates rates; // safe publication
private final CountDownLatch ready = new CountDownLatch(1); // one-shot event
private final Semaphore upstreamSlots = new Semaphore(8); // bulkhead
private final ConcurrentHashMap<String, CompletableFuture<BigDecimal>> cache =
new ConcurrentHashMap<>();
private final Upstream upstream;
private final Executor io;
PriceService(Upstream upstream, Executor io) { this.upstream = upstream; this.io = io; }
public void reloadRates(Map<String, BigDecimal> fresh, long version) {
rates = new Rates(fresh, version); // readers see old or new, never a mix
ready.countDown(); // no effect after the first call
}
public BigDecimal price(String sku, String currency) throws Exception {
if (!ready.await(5, TimeUnit.SECONDS)) throw new IllegalStateException("rates not loaded");
Rates r = rates; // read the reference once
BigDecimal base = cache.computeIfAbsent(sku, this::load).get(2, TimeUnit.SECONDS);
BigDecimal fx = r.perCurrency().get(currency);
if (fx == null) throw new IllegalArgumentException(currency);
return base.multiply(fx);
}
// Called at most once per key while the entry is absent; returns quickly.
private CompletableFuture<BigDecimal> load(String sku) {
CompletableFuture<BigDecimal> f = CompletableFuture.supplyAsync(() -> {
try {
upstreamSlots.acquire();
try { return upstream.fetch(sku); }
finally { upstreamSlots.release(); }
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new CompletionException(e);
}
}, io);
// Do not cache failures: remove the entry so the next caller retries.
// Async so the hook never runs inside computeIfAbsent, even if f already failed.
f.whenCompleteAsync((v, err) -> { if (err != null) cache.remove(sku, f); }, io);
return f;
}
interface Upstream { BigDecimal fetch(String sku); }
}Walk through what each line guarantees. reloadRates builds a new Rates (whose constructor copies the map into an immutable one) and assigns it to a volatile field; that write happens-before any read of rates that sees it, so a reader can never observe a half-built table. price reads rates into a local exactly once, so a reload in the middle of a request cannot mix old and new rates in one answer.
computeIfAbsent on ConcurrentHashMap runs the mapping function at most once per absent key and makes other callers for that key wait for it, which is exactly why the function only creates a future and does no I/O: a slow function would block other updates to the map. The 300 concurrent requests all get the same future and all wait on it, and the future's completion happens-before every get that returns its value. The whenCompleteAsync hook removes failed futures using the two-argument remove, so it cannot delete a newer entry put there by a retry; it runs on the executor because a future that fails instantly would otherwise run the hook inside computeIfAbsent, which the map forbids.
The semaphore is acquired inside the async task, so callers never hold a permit while waiting, and release sits in finally so an exception cannot leak a permit. Every wait in the class has a timeout. This cache has no eviction, which is fine for the demonstration and wrong for production; swap it for a caching library with the same single-flight semantics once the mechanism is clear.
Failure modes
The failures below account for most production concurrency incidents.
- Check-then-act.
if (!map.containsKey(k)) map.put(k, v)on a concurrent map is still a race, because the two calls are individually safe but not atomic together. UseputIfAbsent,computeIfAbsentor a lock around both. - Publishing too early. Starting a thread or registering a listener from inside a constructor lets other threads see
thisbefore construction finishes. Publish only fully built objects. - Lost wakeups and spurious wakeups. Waiting with
ifinstead ofwhile, or signalling without holding the lock that guards the predicate, leaves threads asleep forever or awake on a false condition. - Deadlock. Two locks taken in different orders on two paths. Fix with a global lock order,
tryLockwith a timeout, or by not holding one lock while taking another. - Holding a lock across slow work. I/O, logging to a blocking appender or calling a callback inside a critical section turns a microsecond lock into a 50 ms one; throughput collapses under load and looks like a CPU-idle outage.
- Contention and false sharing. Correct code can still be slow: one hot lock serialises a 64-core machine, and two unrelated counters on the same cache line ping-pong between cores.
- Priority inversion and starvation. A low-priority thread holding a lock blocks a high-priority one; unfair locks can starve a thread indefinitely under steady contention.
Testing and diagnosing concurrent code
Concurrency bugs rarely show up in unit tests, so you need tools that make them more likely or detect them directly.
- Race detectors. Go's
go test -raceand ThreadSanitizer for C, C++ and Rust (-fsanitize=thread) instrument memory accesses and report unordered conflicting pairs with both stack traces. Run them in CI on your concurrent tests; the slowdown is acceptable there. - Stress harnesses. For Java, the OpenJDK jcstress harness runs tiny actors millions of times across threads and reports every observed outcome, including ones the memory model allows but you did not expect.
- Thread dumps.
jstack <pid>orjcmd <pid> Thread.printshows who holds which monitor and who waits for it, and the JVM reports detected Java-level deadlocks at the end of the dump. - Lock profiling. Java Flight Recorder records monitor-blocked and park events with durations; on Linux, off-CPU profiling shows threads parked on futexes. Profile under realistic load, because contention only exists under load.
Trade-offs
Coarse locks are easy to reason about and scale badly; fine-grained locks scale better and multiply deadlock risk. Immutable snapshots make reads free but allocate on every update and suit read-mostly data. Atomics avoid blocking but only protect one variable; the moment an invariant spans two, you need a lock or a single immutable object that holds both. Message passing removes shared mutable state but adds queues that must be bounded and monitored, and turns some bugs into latency instead of corruption. Most good designs mix them: confinement and immutability by default, a few well-named locks for the genuinely shared mutable core, and queues at the boundaries.
Further reading on this site: locks and mutexes in depth, condition variables, semaphores, the Java memory model and deadlock.
What to do next
- List every piece of state your service shares between threads, and for each pick one owner: confined, immutable, atomic, locked or queued.
- For each shared mutable field, write down the happens-before edge that orders its writes and reads; fix any field where you cannot.
- Replace check-then-act sequences on concurrent collections with their atomic methods.
- Add timeouts to every blocking wait and make sure every acquire has a release in
finally. - Move I/O and callbacks out of critical sections.
- Turn on a race detector or a stress harness in CI for code with shared state.
- Profile lock contention under production-like load before adding finer-grained locks.