A race condition is a bug whose outcome depends on the timing of events you do not control: which thread the scheduler runs first, which core sees a write first, which request reaches the database first. The code is correct in every interleaving you imagined while writing it and wrong in one you did not. That is why races survive code review, pass the test suite thousands of times, and then corrupt a balance or double-ship an order once a week in production.

This article treats races as a family rather than a single bug. You will learn the difference between a data race and a race condition, the five shapes that account for nearly every race in real code, how to make each one fail on demand, which tools find them, and how to choose a fix.

Architecture at a glance

Check-then-act: two threads, one seat leftThread Awants the last seatThread Bwants the last seatseats_left = 1shared stateread seats_leftsees 1read seats_leftalso sees 11 > 0, so bookwrites 01 > 0, so bookwrites 0Two bookings, one seatinvariant broken; no crash, no errorthe window: between the checkand the act, the fact can changeTime flows downward. Any interleaving where B reads before A writes produces the bug.
Two threads both pass the check before either acts. Every access may be synchronised, yet the invariant still breaks, because the check and the act are separate steps.

From first principles: data race versus race condition

Two terms get mixed up constantly, and the difference decides which tool can help you.

A data race is a precise, language-level property: two threads access the same memory location concurrently, at least one access is a write, and nothing orders them (no lock, no atomic, no happens-before edge). In C and C++ a data race is undefined behaviour, so the compiler may assume it never happens and optimise accordingly. In Java the program keeps defined, if surprising, semantics: reads can see stale values. Go limits the damage for single-word values, but races on interfaces, slices, strings or maps can crash the program or corrupt memory. Data races are mechanical, which is why tools can find them automatically.

A race condition is a semantic property: the program's correctness depends on timing. You can have a race condition with zero data races. Wrap the read and the write of the seat counter in the diagram in two separate locked methods and every access is synchronised, yet two bookings still succeed, because the invariant needed the read and the write to be one atomic step. Conversely, a data race on a statistics counter you only ever eyeball may not be a correctness problem at all, though in C++ it is still undefined behaviour and must be fixed.

The working definition to keep: a race condition exists when an invariant must hold across several steps, and other threads can observe or change state between those steps. Every fix either removes the sharing, makes the steps one atomic step, or makes the later step re-check the earlier one.

The five shapes of a race

Almost every race in production code is one of five shapes. Learning to name them is most of the work, because once you see the shape the fix follows.

ShapeWhat it looks likeTypical fix
Read-modify-writecount++, balance = balance - amountAtomic operation, or a lock around read and write
Check-then-actif not exists: create, if (map.get(k) == null) map.put(k, v)One atomic primitive (putIfAbsent, unique constraint, CAS)
Time-of-check to time-of-use (TOCTOU)check a file, a permission or a row, then use it laterOperate on a handle, not a name; check and use in one step
Unsafe publicationlazy init, double-checked locking without volatile, sharing a half-built objectImmutable objects, safe publication, sync.Once or a holder class
Compound operations on safe partstwo calls on a thread-safe collection, or two services each correct aloneSingle atomic method, transaction, or ownership by one thread

Read-modify-write is the lost update: two threads read 5, both write 6, one increment vanishes. It is covered with hardware detail in atomics and compare-and-swap, so this page focuses on the other four, which are more common in application code and harder to spot.

The shapes in code

Check-then-act on a thread-safe collection. Using a concurrent map makes each call safe, not your sequence of calls. The classic symptom is duplicated work or a lost entry under load:

// Java: check-then-act on a thread-safe map is still a race
ConcurrentHashMap<String, Session> sessions = new ConcurrentHashMap<>();

// BROKEN: two threads can both see null and both create a session
Session getSession(String user) {
    Session s = sessions.get(user);          // check
    if (s == null) {
        s = new Session(user);               // expensive, has side effects
        sessions.put(user, s);               // act: second put overwrites the first
    }
    return s;
}

// FIXED: one atomic operation; the function runs at most once per key
Session getSessionFixed(String user) {
    return sessions.computeIfAbsent(user, Session::new);
}

TOCTOU. The check and the use refer to the same name, but the thing behind the name can change in between. It is the same bug at the filesystem, permission and database level, and at the filesystem level it is a security hole, because an attacker can win the race on purpose:

# Python: TOCTOU on the filesystem
import os

# BROKEN: between exists() and open(), another process can create the file,
# or swap it for a symlink to something you should not overwrite
if not os.path.exists(path):
    with open(path, "w") as f:
        f.write(data)

# FIXED: ask the kernel to check and create in one system call
fd = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
with os.fdopen(fd, "w") as f:
    f.write(data)

Unsafe publication. One thread builds an object and stores a reference; another reads the reference and uses the object. Without a happens-before edge, the reader can see the reference before it sees the fields, or two threads can both initialise. The fix is a primitive designed for one-time initialisation, or an immutable object published through a properly synchronised field. The memory model article explains why the hardware and compiler are allowed to reorder these writes.

// Go: unsafe lazy initialisation, and the fix
var cfg *Config

func GetConfigBroken() *Config {
    if cfg == nil {          // data race: unsynchronised read
        cfg = loadConfig()   // data race: unsynchronised write; may run twice
    }
    return cfg
}

var (
    once    sync.Once
    cfgSafe *Config
)

func GetConfig() *Config {
    once.Do(func() { cfgSafe = loadConfig() }) // runs exactly once, publishes safely
    return cfgSafe
}

Compound operations across components. The same shape appears between services: a service reads stock from one API, then reserves it with another. Each call is correct; the pair is not. At that scale locks do not exist, so the fix is a conditional write in the store that owns the data, such as a SQL UPDATE ... WHERE stock >= 1 that reports how many rows it changed, or a compare-and-set on a version column.

Worked example: the double-click withdrawal

A wallet service lets users withdraw from a balance. Support reports that a handful of accounts went negative during a promotion, always for users who double-clicked. The code checks the balance and then subtracts, and it runs on Python, where the global interpreter lock makes developers assume this is safe. It is not: the GIL makes individual bytecodes atomic, not your two lines, and a thread switch can happen between the check and the subtraction.

The first step is to make the race fail every time instead of once in ten thousand runs. Random sleeps are unreliable; a barrier placed inside the window forces both threads to finish the check before either acts. The test hook costs nothing in production and turns a flaky bug into a deterministic test:

# Python: a race that the GIL does not prevent, made deterministic
import threading

class Account:
    def __init__(self, balance):
        self.balance = balance
        self.lock = threading.Lock()

    def withdraw_broken(self, amount, gate=None):
        if self.balance >= amount:          # check
            if gate:                        # test hook: widen the window
                gate.wait()
            self.balance -= amount          # act
            return True
        return False

    def withdraw(self, amount):
        with self.lock:                     # check and act are one critical section
            if self.balance >= amount:
                self.balance -= amount
                return True
            return False

acct = Account(100)
gate = threading.Barrier(2)                 # both threads pass the check, then proceed
results = []
ts = [threading.Thread(target=lambda: results.append(acct.withdraw_broken(80, gate)))
      for _ in range(2)]
for t in ts: t.start()
for t in ts: t.join()
print(results, acct.balance)                # [True, True] -60, every single run

With the hook, the broken version returns two successes and a balance of -60 on every run. The fixed version holds the lock across the check and the subtraction, so the second thread sees 20 and refuses. The production fix went one step further, because the balance actually lived in a database shared by several processes and a thread lock protects only one process: the withdrawal became UPDATE accounts SET balance = balance - :amt WHERE id = :id AND balance >= :amt, and the service treats zero updated rows as insufficient funds. That moves the atomic step to the component that owns the data, which is the general rule.

Finding races: detectors and stress tests

Detection tools fall into two groups, and they find different things.

Dynamic data-race detectors instrument memory accesses and track happens-before relations at run time. ThreadSanitizer powers the C, C++ and Go detectors; the Go race detector is built on it. They report real data races with both stack traces, with no false positives in the common case, but only on code paths the run actually executes. The cost is substantial, roughly several times slower and several times more memory according to the tools' own documentation, so run them in CI and in a canary, not across the fleet.

Stress and interleaving explorers such as jcstress for the JVM run tiny tests millions of times across cores and report which outcomes occurred. They find races in lock-free code and memory-ordering assumptions that a detector would accept.

# C/C++: ThreadSanitizer (clang or gcc)
clang++ -fsanitize=thread -g -O1 server.cpp -o server && ./server

# Go: the race detector, in tests and in a canary binary
go test -race ./...
go build -race -o app-race ./cmd/app

// Java: jcstress explores interleavings of small actors
@JCStressTest
@Outcome(id = "1, 1", expect = Expect.ACCEPTABLE_INTERESTING, desc = "both saw the seat")
@Outcome(expect = Expect.ACCEPTABLE, desc = "only one booking")
@State
public class SeatRace {
    int seats = 1;
    @Actor public void a(II_Result r) { if (seats > 0) { seats = 0; r.r1 = 1; } }
    @Actor public void b(II_Result r) { if (seats > 0) { seats = 0; r.r2 = 1; } }
}

Neither group finds race conditions that contain no data race. The seat booking with two separately locked methods is invisible to ThreadSanitizer, because every access is synchronised. For those you need invariant checks: assertions that fire when the balance goes negative, database constraints that reject the second booking, and deterministic tests like the barrier test above. Safe Rust removes data races at compile time but likewise cannot stop a check-then-act across two lock acquisitions.

Operating concurrent code

  • Treat every race report as a bug. A data race that looks benign today becomes a broken invariant after the next refactor or compiler upgrade. Keep the race detector job green, not merely monitored.
  • Look for the symptoms. Duplicates, counters that drift from the source of truth, rare null dereferences on lazily initialised fields, and bugs that disappear when you add logging (the log call changes timing) are race signatures.
  • Push invariants down to the owner. Unique indexes, conditional updates and idempotency keys enforce invariants even when several processes race; in-process locks cannot.
  • Prefer ownership to sharing. State owned by one thread or actor, with messages in and out, cannot race. Immutable values can be shared freely.

Failure modes

  • Fixing the symptom with a sleep. Adding a delay changes timing and hides the race until the load profile changes.
  • Locking each step separately. Two correctly locked methods called in sequence still race; the fix must make the sequence atomic.
  • Double-checked locking without a happens-before edge. Readers can see a half-constructed object. Use the language's once primitive or a holder idiom.
  • Trusting a runtime lock to cover multiple processes. Two pods each hold their own mutex and both proceed.
  • Over-correcting into deadlock. Wrapping everything in nested locks trades a race for a lock-ordering bug; see the deadlock guide.

Trade-offs

Every fix trades something. A coarse lock is easy to reason about and limits throughput. Atomics are fast but cover only one variable, so they cannot protect an invariant that spans two. Conditional writes in the database survive multiple processes and crashes, and cost a round trip and a retry path. Ownership by a single thread removes the race class entirely but adds a queue, and the queue adds latency and back-pressure concerns. Choose the narrowest mechanism that covers the whole invariant; when in doubt, pick the one that is easier to prove correct and measure before optimising. The locks and mutexes guide covers granularity and contention in detail, and structured concurrency shows how to limit how much state tasks share in the first place.

What to do next

  1. Run your test suite under a race detector this week: go test -race, a ThreadSanitizer build, or jcstress for concurrent Java classes.
  2. Grep for the five shapes: get-then-put on maps, exists-then-open on files, if x == null lazy initialisation, read-then-update on shared counters, and read-then-write across service calls.
  3. For each invariant that matters (no negative balance, one booking per seat, one session per user), write down which component enforces it atomically. If the answer is nobody, add a constraint or conditional write.
  4. Turn the next race you find into a deterministic test with a barrier or latch inside the window, and keep it as a regression test.
  5. Add a CI job that keeps the race detector green and treat new reports as release blockers.
Key takeaway: A race condition exists whenever an invariant spans several steps and another thread or process can act between them. Data races are the mechanical subset that ThreadSanitizer, the Go race detector and jcstress can find; race conditions without data races need invariant checks, constraints and deterministic tests. Learn the five shapes, read-modify-write, check-then-act, TOCTOU, unsafe publication and compound operations, and fix each by making the whole invariant one atomic step in the component that owns the data, or by removing the sharing.