A mutex answers one question: may I touch this data now? Many concurrent programs need a different one: may I proceed yet? A consumer needs the queue to be non-empty, a worker needs a job, a shutdown needs every task to finish. Spinning wastes a core; sleeping for a fixed time adds latency. A condition variable lets a thread sleep until another thread changes the state it is waiting on, and then wake up holding the lock.

This page covers the three parts every correct use has, the one loop shape that is correct in every language, the two classic bugs, a worked bounded buffer, signal versus broadcast, how runtimes build it on the kernel, and when to use something else.

Advertisement

Three parts: mutex, predicate, wait queue

A condition variable is not a condition. It holds no value and remembers nothing. It is a queue of sleeping threads with two operations: wait (join the queue) and notify (wake one or all of the queue). The condition itself is a predicate over shared state, such as the queue holding at least one item, and that state is protected by a mutex. Every correct use involves all three, and the rule that ties them together is simple: read and write the predicate's state only while holding the mutex.

The key operation is wait. Called while holding the mutex, it releases the mutex and puts the thread to sleep as one atomic step, then, when the thread is woken, reacquires the mutex before returning. Atomicity is the whole point. If releasing and sleeping were two separate steps, another thread could change the state and notify in the gap, and the notification would reach nobody.

A condition variable is a wait queue tied to a mutex and a predicateMutexprotects the shared stateShared statepredicate: items > 0Wait queuethreads parked in wait()Signallerchanges state, then notifywait: release mutex, parknotify: wake one / allwoken: reacquire mutex, re-check predicateunder mutex: modifyReleasing the mutex and parking happen atomically, so no notify can slip between them
The mutex protects the state, the predicate is a test on that state, and the condition variable is only the queue of threads waiting for the predicate to become true.

The canonical pattern in four APIs

Every language spells it the same way: lock, loop while the predicate is false, wait, act, unlock. In Java with java.util.concurrent.locks:

final Lock lock = new ReentrantLock();
final Condition notEmpty = lock.newCondition();
final Deque<Task> queue = new ArrayDeque<>();

Task take() throws InterruptedException {
  lock.lock();
  try {
    while (queue.isEmpty()) {          // re-check after every wakeup
      notEmpty.await();                // releases lock, parks, reacquires
    }
    return queue.removeFirst();
  } finally {
    lock.unlock();
  }
}

void put(Task t) {
  lock.lock();
  try {
    queue.addLast(t);
    notEmpty.signal();                 // must hold the lock in Java
  } finally {
    lock.unlock();
  }
}

In C++, std::condition_variable works with std::unique_lock<std::mutex>, and the overload that takes a predicate writes the loop for you:

std::mutex m;
std::condition_variable not_empty;
std::deque<Task> queue;

Task take() {
    std::unique_lock<std::mutex> lk(m);
    not_empty.wait(lk, [&] { return !queue.empty(); });   // loop is inside wait()
    Task t = std::move(queue.front());
    queue.pop_front();
    return t;
}

void put(Task t) {
    {
        std::lock_guard<std::mutex> lk(m);
        queue.push_back(std::move(t));
    }
    not_empty.notify_one();            // allowed after unlock in C++ and POSIX
}

In C with POSIX threads:

pthread_mutex_lock(&m);
while (count == 0)
    pthread_cond_wait(&not_empty, &m);   /* atomically unlocks m and sleeps */
item = items[--count];
pthread_mutex_unlock(&m);

Java's built-in monitors give every object one implicit condition through wait(), notify() and notifyAll(), used inside synchronized. The explicit Condition API is preferable when you need more than one wait queue per lock, timed waits that report the remaining time, or a fair lock. Python's threading.Condition, Go's sync.Cond and Rust's std::sync::Condvar follow the same model.

One difference matters. Java requires holding the lock to signal and throws IllegalMonitorStateException otherwise, and Python raises RuntimeError. POSIX and C++ allow notifying after unlocking, which can save the woken thread from immediately blocking on a mutex the signaller still holds.

Advertisement

Why while, never if

Almost every condition variable in use has Mesa semantics, named after Xerox's Mesa language, whose monitors Lampson and Redell described in 1980. A notify is a hint that the state may have changed, not a guarantee that the predicate is true when the woken thread runs. Between the notify and the moment the woken thread reacquires the mutex, any other thread can take the mutex first and change the state. With two consumers and one item, the second consumer to run finds the queue empty again. Code that waits under if then removes from an empty queue.

Hoare semantics, the alternative, hands the mutex straight from signaller to waiter so the predicate holds on wakeup; it costs extra context switches and mainstream runtimes do not offer it.

There is a second reason. POSIX, C++ and Java's Object.wait all explicitly permit spurious wakeups: a waiting thread may return without any notify at all. Implementations allow this because guaranteeing otherwise would cost performance, for example when a signal interrupts the underlying kernel wait. The loop handles both cases with the same three lines, so write it every time, even where you believe it cannot matter.

Lost wakeups

The opposite bug is a notify that nobody receives, leaving a thread asleep forever. It happens when the state behind the predicate is changed without holding the mutex:

// BROKEN: the flag is written without the lock.
volatile boolean ready = false;

// waiter                                  // signaller
lock.lock();                               //
while (!ready)                             //   reads ready == false ...
                                           ready = true;          // no lock held
                                           cond.signal();         // nobody waiting yet: lost
    cond.await();                          // ... sleeps forever
lock.unlock();

The waiter checked the flag and was about to wait, the signaller set the flag and signalled into an empty queue, and then the waiter went to sleep on a state that was already true. Marking the flag volatile fixes visibility, not this race. Only the mutex closes the window, because the waiter holds it from the check until wait atomically releases it, and the signaller cannot change the flag in between.

The rules are short: change the state under the mutex; check the predicate under the same mutex; never wait without checking first, since a notify that happened before you started waiting is not stored anywhere. A condition variable has no memory; a semaphore does, which is one reason to choose a semaphore when you really want to count events.

Worked example: a bounded buffer

A bounded buffer needs two predicates: producers wait for not full, consumers wait for not empty. Both are about the same state, so both conditions share one mutex. Python's threading.Condition accepts an existing lock, which makes this explicit:

import threading
from collections import deque

class BoundedBuffer:
    def __init__(self, capacity):
        self.items = deque()
        self.capacity = capacity
        self.lock = threading.Lock()
        self.not_full = threading.Condition(self.lock)    # two conditions,
        self.not_empty = threading.Condition(self.lock)   # one mutex

    def put(self, item):
        with self.lock:
            while len(self.items) >= self.capacity:      # while, never if
                self.not_full.wait()
            self.items.append(item)
            self.not_empty.notify()                       # one item -> one getter

    def get(self):
        with self.lock:
            while not self.items:
                self.not_empty.wait()
            item = self.items.popleft()
            self.not_full.notify()                        # one slot -> one putter
            return item

buf = BoundedBuffer(capacity=8)
N_PRODUCERS, PER_PRODUCER, N_CONSUMERS = 4, 25_000, 4
STOP = object()
totals = []

def producer(base):
    for i in range(PER_PRODUCER):
        buf.put(base + i)

def consumer():
    s = 0
    while (item := buf.get()) is not STOP:
        s += item
    totals.append(s)

consumers = [threading.Thread(target=consumer) for _ in range(N_CONSUMERS)]
producers = [threading.Thread(target=producer, args=(p * PER_PRODUCER,)) for p in range(N_PRODUCERS)]
for t in consumers + producers:
    t.start()
for t in producers:
    t.join()
for _ in consumers:
    buf.put(STOP)                     # one poison pill per consumer
for t in consumers:
    t.join()

n = N_PRODUCERS * PER_PRODUCER
print(f"consumed sum {sum(totals):,}; expected {n * (n - 1) // 2:,}")
print(f"buffer empty at exit: {not buf.items}")

Running it prints:

consumed sum 4,999,950,000; expected 4,999,950,000
buffer empty at exit: True

Four producers push 100,000 numbers through an eight-slot buffer to four consumers, and the sum matches exactly: nothing lost or delivered twice. Two conditions mean a put wakes only getters and a get wakes only putters; with one shared condition, a putter's notify could wake another putter, which finds the buffer full and sleeps again while the consumer keeps waiting. Each operation calls notify() because one item or one slot serves exactly one waiter. One sentinel per consumer stops them only after every real item is taken.

Signal or broadcast

Wake one thread when any single waiter can make use of the change and each change helps exactly one waiter: one item in, one consumer out. Wake all when the change may let several waiters proceed, when waiters wait for different predicates on the same condition, or when you cannot tell which waiter can proceed. Examples are a shutdown flag, a barrier opening, a configuration reload, or releasing a reader-writer lock to all readers.

Broadcast is always correct if the wait loop is correct, because a woken thread whose predicate is false just waits again. It costs a thundering herd: with 64 waiters, all 64 wake, contend for the mutex, and 63 go back to sleep. Signal is cheaper but correct only under the conditions above, and the classic bug is a single condition shared by two kinds of waiter where signal wakes the wrong kind and the right one sleeps forever. When in doubt, use separate conditions per predicate and signal; if you cannot separate them, broadcast.

How it is built

On Linux, the building block is the futex system call: a thread can ask the kernel to sleep only if a given memory word still holds an expected value, and another thread can wake sleepers on that word. That check-and-sleep happens atomically in the kernel, which is what lets user-space code release a mutex and then sleep without a lost-wakeup window. When no one is waiting, signal and wait need no system call at all.

glibc's pthread_cond_t is built on futexes; glibc 2.25 rewrote it because the old algorithm could let a signal be consumed by a thread that started waiting after the signal was sent.

In Java, ReentrantLock's conditions are implemented by AbstractQueuedSynchronizer. await() adds the thread to the condition's own queue, fully releases the lock (saving the hold count, so reentrant holds are restored later), and parks. signal() does not wake the thread directly: it moves the first waiter from the condition queue to the lock's queue, where it competes to reacquire the lock like any other thread. That is Mesa semantics, visible in the data structure.

Waking a thread is not free: it must be scheduled, often on a core with cold caches. When waits are very short and contended, brief spinning or a lock-free queue can win.

Timed waits and clocks

Real systems rarely wait forever. A timed wait returns after a deadline even if nothing was signalled, and the loop must handle that:

long nanos = TimeUnit.SECONDS.toNanos(5);
lock.lock();
try {
  while (queue.isEmpty()) {
    if (nanos <= 0L) return null;           // deadline passed: give up
    nanos = notEmpty.awaitNanos(nanos);     // returns the time still left
  }
  return queue.removeFirst();
} finally {
  lock.unlock();
}

Compute the deadline once and pass the remaining time on each iteration; restarting a fixed timeout after every spurious or useless wakeup can wait forever. Use a monotonic clock: pthread_cond_timedwait uses the realtime clock unless you set CLOCK_MONOTONIC with pthread_condattr_setclock, and a wall-clock jump from an NTP step would then stretch or cut the wait. C++'s wait_for is specified against a steady clock, and Java's awaitNanos uses relative nanoseconds.

When to use something else

Condition variables are the low-level tool. Prefer a higher-level primitive when one fits. A blocking queue (Java's ArrayBlockingQueue, Python's queue.Queue) is the bounded buffer above, written and tested once; both are built on condition variables. Channels express hand-offs between threads more directly, and Go's documentation for sync.Cond points users to channels for most cases. A semaphore counts permits and remembers releases. A latch or barrier expresses wait for N events. Use a raw condition variable to build those, or for compound predicates no library structure holds.

Failure modes

  • Waiting under if. A woken thread acts on a predicate that another thread has already made false. Always loop.
  • Lost wakeup. State changed without the mutex, or wait called without first checking the predicate. The thread sleeps forever.
  • Wrong waiter woken. One condition shared by different predicates, plus signal: the notify wakes a thread that cannot proceed.
  • Waiting while holding another lock. Wait releases only its own mutex. Any other lock the thread holds stays held while it sleeps, which invites deadlock.
  • Priority inversion. A high-priority waiter depends on a low-priority thread to signal; see priority inversion.
  • Destroying a condition too early. In C and C++, notifying after unlock can race with a woken thread destroying the object that owns the condition variable. Signal under the lock when the waiter may free the state.
  • Ignoring interruption. Java's await() throws InterruptedException; swallowing it breaks cancellation. Restore the interrupt flag or propagate it.

Trade-offs

ChoiceGainCost
Condition variable vs spinningNo CPU burned while waitingContext switches and wake latency
Signal (notify one)Wakes only what can make progressCorrect only for interchangeable waiters
Broadcast (notify all)Always correct with a proper loopThundering herd on the mutex
Several conditions per lockTargeted wakeupsMore objects to keep consistent
Notify after unlockWoken thread need not block on the mutexDestruction races in C and C++
Blocking queue or channel insteadCorrectness written onceLess flexible predicates

What to do next

  1. Search your codebase for every wait call and check that each sits inside a while loop over a predicate.
  2. For each condition, write down its predicate and the mutex that protects that state; any write to the state outside that mutex is a bug.
  3. Replace hand-written producer-consumer code with your platform's blocking queue or channel where it fits.
  4. Split any condition shared by two kinds of waiter into one condition per predicate.
  5. Audit timed waits: compute one deadline, loop on the remaining time, and use a monotonic clock.
  6. Run the bounded buffer above with more threads and a capacity of 1 to see the protocol hold under contention.
  7. Read the memory model page to see why the mutex also makes the state change visible to the woken thread.
Key takeaway: A condition variable is a wait queue tied to a mutex and a predicate over shared state. Change and check the state only under the mutex, and always wait inside a while loop, because notifies are hints under Mesa semantics and spurious wakeups are allowed. A condition remembers nothing, so a notify with no waiter is lost. Use one condition per predicate and signal when any single waiter can use the change; broadcast otherwise. Prefer blocking queues, channels or semaphores when they fit, and use monotonic deadlines for timed waits.