Two threads that never share anything need no synchronization at all. The moment they share a counter, a cache, a queue or a flag, you need two guarantees that hardware and compilers do not give you by default: that one thread's half-finished update is never seen or overwritten by another, and that when one thread finishes something, the other actually sees it. Locks, condition variables, semaphores, latches, atomics and channels are all different ways of buying those two guarantees, and each fits a different shape of problem.

This article is the map rather than a tour of one primitive. It explains the two jobs synchronization does, the single rule (happens-before) that every primitive is built on, how to pick a primitive from the shape of your coordination problem, and then builds a small but realistic Java service that uses four different primitives, each for the job it is best at. It ends with failure modes, how to test concurrent code and a checklist.

The map

Thread Awrites shared stateThread Breads shared stateShared statefields, maps, queuesreleaseacquireExclusionmutex, rwlock, spinlockWaiting for a conditioncondvar, monitorCountingsemaphore, rate limitOne-shot eventslatch, future, onceSingle-word updatesatomics, CASHanding over dataqueue, channelevery primitive creates ahappens-before edge fromrelease in A to acquire in B
Shared state is safe when every write in one thread is ordered before the reads in another by some primitive's release-acquire edge; the primitives differ in the coordination shape they fit.

Two jobs: exclusion and visibility

Synchronization does two separate jobs, and most bugs come from solving one and forgetting the other.

Exclusion makes a group of operations atomic with respect to other threads. The classic example is count = count + 1: it is a load, an add and a store, and two threads can both load 41 and both store 42. A transfer between two accounts, a check-then-insert into a map, and an update of two fields that must agree are the same problem in larger form. Exclusion says: while I am in this region, nobody else touching the same data is.

Ordering and visibility make one thread's writes visible to another, in an order it can rely on. Modern CPUs keep writes in store buffers and caches, and compilers keep values in registers and reorder independent instructions. Without synchronization, thread B can spin forever on a ready flag that thread A set long ago, or see ready == true but a stale value in the data the flag was guarding. No interleaving explains that result; it comes from reordering, which is why race bugs are so hard to reason about by staring at source code.

A third job, waiting, sits on top of these: a thread needs to sleep until some condition holds (the queue is non-empty, five workers have finished, a permit is free) without burning a CPU core and without missing the moment the condition becomes true.

Happens-before: the one rule underneath

Every mainstream memory model (Java, C++11 and later, Go, Rust through C++'s) is defined in terms of one relation: happens-before. If write W happens-before read R, R is guaranteed to see W or something later. If two accesses to the same location are not ordered by happens-before and at least one is a write, you have a data race. In C and C++ a data race is undefined behaviour; in Java the program stays memory-safe, but the racing read may see stale or partially ordered values.

Within one thread, program order gives you happens-before for free. Between threads you only get it from synchronization: a mutex unlock happens-before the next lock of the same mutex; a write to a Java volatile field happens-before every later read that sees it; Thread.start happens-before the started thread's first action; completing a future happens-before a get that returns its value; putting an item in a BlockingQueue happens-before taking it out. The relation is transitive, which is what makes safe publication work: build an object fully, then publish its reference through any of these edges, and the reader sees the fully built object.

This gives a test you can apply to any piece of shared state: for every pair of accesses from different threads where one is a write, name the happens-before edge that orders them. If you cannot name it, you have a race, however unlikely it looks.

The toolbox, sorted by question

Primitives are best sorted by the question they answer, not by how they are implemented.

QuestionPrimitiveTypical Java formWatch out for
Only one thread may run this regionMutexsynchronized, ReentrantLockHolding it across I/O; lock ordering
Many readers, rare writersRead-write lockReentrantReadWriteLock, StampedLockWriter starvation; often slower than a mutex for short sections
Sleep until a predicate is trueCondition variablewait/notifyAll, ConditionMust re-check the predicate in a loop
At most N at onceCounting semaphoreSemaphoreLeaked permits when release is skipped
Wait for one event, onceLatch, future, once-initCountDownLatch, CompletableFutureWaiting forever if the event never fires
N threads meet, then all continueBarrierCyclicBarrier, PhaserOne dead party blocks everyone
Update one variable without a lockAtomic, compare-and-swapAtomicLong, LongAdderCompound invariants across two variables
Pass data, not share itQueue, channelArrayBlockingQueueUnbounded queues hide overload
Swap a whole consistent viewImmutable snapshot + volatile refvolatile, AtomicReferenceMutating the snapshot after publishing it

Choosing a primitive

A short decision procedure covers most code:

  1. Can you avoid sharing? Confine state to one thread (an actor, an event loop, a per-request object) or make it immutable. Immutable data needs only safe publication, never locking.
  2. Is the shared thing a single value? Use an atomic. Counters with heavy write contention do better with LongAdder, which spreads updates over cells and sums them on read.
  3. Is it a consistent multi-field view that changes rarely? Build a new immutable object and swap the reference. Readers pay one volatile read and never block.
  4. Is it producer-consumer? Use a bounded queue or channel and let it carry both the data and the happens-before edge.
  5. Do several fields change together often? Use a mutex, keep the critical section short, and never call unknown code while holding it.
  6. Do threads need to wait for a state, a count or an event? Use the matching higher-level tool (latch, semaphore, future, barrier) before writing a condition variable by hand.

The order matters. Each step down the list adds a way to deadlock, lose a wakeup or contend. Lock-free algorithms built from many compare-and-swap steps are deliberately missing from the list: they are worth it inside libraries, rarely in application code.

Worked example: a price service with four primitives

Consider a pricing service. It answers price(sku) from a cache, loads missing prices from a slow upstream API, applies a currency table that an operator can reload at any time, and must not answer before its first currency table is loaded. Four requirements, four primitives:

  • The currency table is a consistent multi-field view that changes rarely: an immutable snapshot behind a volatile reference.
  • When 300 requests for the same uncached SKU arrive together, the upstream must be called once, not 300 times: single-flight loading with computeIfAbsent and a CompletableFuture.
  • The upstream tolerates at most 8 concurrent calls: a semaphore as a bulkhead.
  • Requests must wait until the first table arrives: a latch.
import java.math.BigDecimal;
import java.util.Map;
import java.util.concurrent.*;

public final class PriceService {
    // Immutable view: built fully, then published by one volatile write.
    record Rates(Map<String, BigDecimal> perCurrency, long version) {
        Rates { perCurrency = Map.copyOf(perCurrency); }
    }

    private volatile Rates rates;                                   // safe publication
    private final CountDownLatch ready = new CountDownLatch(1);     // one-shot event
    private final Semaphore upstreamSlots = new Semaphore(8);       // bulkhead
    private final ConcurrentHashMap<String, CompletableFuture<BigDecimal>> cache =
            new ConcurrentHashMap<>();
    private final Upstream upstream;
    private final Executor io;

    PriceService(Upstream upstream, Executor io) { this.upstream = upstream; this.io = io; }

    public void reloadRates(Map<String, BigDecimal> fresh, long version) {
        rates = new Rates(fresh, version);   // readers see old or new, never a mix
        ready.countDown();                   // no effect after the first call
    }

    public BigDecimal price(String sku, String currency) throws Exception {
        if (!ready.await(5, TimeUnit.SECONDS)) throw new IllegalStateException("rates not loaded");
        Rates r = rates;                                  // read the reference once
        BigDecimal base = cache.computeIfAbsent(sku, this::load).get(2, TimeUnit.SECONDS);
        BigDecimal fx = r.perCurrency().get(currency);
        if (fx == null) throw new IllegalArgumentException(currency);
        return base.multiply(fx);
    }

    // Called at most once per key while the entry is absent; returns quickly.
    private CompletableFuture<BigDecimal> load(String sku) {
        CompletableFuture<BigDecimal> f = CompletableFuture.supplyAsync(() -> {
            try {
                upstreamSlots.acquire();
                try { return upstream.fetch(sku); }
                finally { upstreamSlots.release(); }
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
                throw new CompletionException(e);
            }
        }, io);
        // Do not cache failures: remove the entry so the next caller retries.
        // Async so the hook never runs inside computeIfAbsent, even if f already failed.
        f.whenCompleteAsync((v, err) -> { if (err != null) cache.remove(sku, f); }, io);
        return f;
    }

    interface Upstream { BigDecimal fetch(String sku); }
}

Walk through what each line guarantees. reloadRates builds a new Rates (whose constructor copies the map into an immutable one) and assigns it to a volatile field; that write happens-before any read of rates that sees it, so a reader can never observe a half-built table. price reads rates into a local exactly once, so a reload in the middle of a request cannot mix old and new rates in one answer.

computeIfAbsent on ConcurrentHashMap runs the mapping function at most once per absent key and makes other callers for that key wait for it, which is exactly why the function only creates a future and does no I/O: a slow function would block other updates to the map. The 300 concurrent requests all get the same future and all wait on it, and the future's completion happens-before every get that returns its value. The whenCompleteAsync hook removes failed futures using the two-argument remove, so it cannot delete a newer entry put there by a retry; it runs on the executor because a future that fails instantly would otherwise run the hook inside computeIfAbsent, which the map forbids.

The semaphore is acquired inside the async task, so callers never hold a permit while waiting, and release sits in finally so an exception cannot leak a permit. Every wait in the class has a timeout. This cache has no eviction, which is fine for the demonstration and wrong for production; swap it for a caching library with the same single-flight semantics once the mechanism is clear.

Failure modes

The failures below account for most production concurrency incidents.

  • Check-then-act. if (!map.containsKey(k)) map.put(k, v) on a concurrent map is still a race, because the two calls are individually safe but not atomic together. Use putIfAbsent, computeIfAbsent or a lock around both.
  • Publishing too early. Starting a thread or registering a listener from inside a constructor lets other threads see this before construction finishes. Publish only fully built objects.
  • Lost wakeups and spurious wakeups. Waiting with if instead of while, or signalling without holding the lock that guards the predicate, leaves threads asleep forever or awake on a false condition.
  • Deadlock. Two locks taken in different orders on two paths. Fix with a global lock order, tryLock with a timeout, or by not holding one lock while taking another.
  • Holding a lock across slow work. I/O, logging to a blocking appender or calling a callback inside a critical section turns a microsecond lock into a 50 ms one; throughput collapses under load and looks like a CPU-idle outage.
  • Contention and false sharing. Correct code can still be slow: one hot lock serialises a 64-core machine, and two unrelated counters on the same cache line ping-pong between cores.
  • Priority inversion and starvation. A low-priority thread holding a lock blocks a high-priority one; unfair locks can starve a thread indefinitely under steady contention.

Testing and diagnosing concurrent code

Concurrency bugs rarely show up in unit tests, so you need tools that make them more likely or detect them directly.

  • Race detectors. Go's go test -race and ThreadSanitizer for C, C++ and Rust (-fsanitize=thread) instrument memory accesses and report unordered conflicting pairs with both stack traces. Run them in CI on your concurrent tests; the slowdown is acceptable there.
  • Stress harnesses. For Java, the OpenJDK jcstress harness runs tiny actors millions of times across threads and reports every observed outcome, including ones the memory model allows but you did not expect.
  • Thread dumps. jstack <pid> or jcmd <pid> Thread.print shows who holds which monitor and who waits for it, and the JVM reports detected Java-level deadlocks at the end of the dump.
  • Lock profiling. Java Flight Recorder records monitor-blocked and park events with durations; on Linux, off-CPU profiling shows threads parked on futexes. Profile under realistic load, because contention only exists under load.

Trade-offs

Coarse locks are easy to reason about and scale badly; fine-grained locks scale better and multiply deadlock risk. Immutable snapshots make reads free but allocate on every update and suit read-mostly data. Atomics avoid blocking but only protect one variable; the moment an invariant spans two, you need a lock or a single immutable object that holds both. Message passing removes shared mutable state but adds queues that must be bounded and monitored, and turns some bugs into latency instead of corruption. Most good designs mix them: confinement and immutability by default, a few well-named locks for the genuinely shared mutable core, and queues at the boundaries.

Further reading on this site: locks and mutexes in depth, condition variables, semaphores, the Java memory model and deadlock.

What to do next

  1. List every piece of state your service shares between threads, and for each pick one owner: confined, immutable, atomic, locked or queued.
  2. For each shared mutable field, write down the happens-before edge that orders its writes and reads; fix any field where you cannot.
  3. Replace check-then-act sequences on concurrent collections with their atomic methods.
  4. Add timeouts to every blocking wait and make sure every acquire has a release in finally.
  5. Move I/O and callbacks out of critical sections.
  6. Turn on a race detector or a stress harness in CI for code with shared state.
  7. Profile lock contention under production-like load before adding finer-grained locks.
Key takeaway: Synchronization does two jobs: exclusion, so compound updates are atomic, and ordering, so one thread's writes become visible to another. Every primitive provides both by creating a happens-before edge from a release in one thread to an acquire in another, and any shared access you cannot pin to such an edge is a race. Prefer confinement and immutability, then atomics, snapshots and queues, then locks, and reach for latches, semaphores and futures before hand-written condition waits. Keep critical sections short, put timeouts on waits, and test with race detectors and stress harnesses because ordinary tests rarely expose races.