Lock contention is what happens when threads spend time waiting to acquire a lock that another thread holds. It is one of the most common reasons a service stops scaling: you add cores or threads and throughput stays flat or falls, latency rises, and CPU utilisation sits suspiciously below what the load should produce. It is also easy to misdiagnose, because a waiting thread does not show up in an ordinary CPU profile.

This article gives a repeatable method: recognise the symptoms, quantify how busy a lock is, find the lock with the right tool for your runtime, establish whether the cause is long hold times or high acquisition rates, and pick a fix that matches. It includes the queueing arithmetic, worked examples in Java and Go, and a portable instrumented lock you can drop into code you control.

Recognising contention

Contention has a recognisable signature, and the signature matters because the remedies for CPU saturation and for contention are opposite: adding threads helps one and hurts the other.

  • Throughput plateaus or drops as you add threads or load, while total CPU stays well below the number of cores.
  • Tail latency grows much faster than median latency, because a few requests queue behind many holders.
  • Context switches per second climb sharply, visible in vmstat 1 under the cs column or in pidstat -w.
  • System time rises, and syscall tracing shows many futex calls, the Linux primitive on which pthread mutexes and Java thread parking sleep. Go parks goroutines in its own scheduler instead, so in Go the runtime profiles below matter more than syscall counts.
  • Thread dumps taken a few seconds apart show many threads blocked on the same monitor or parked in the same lock class.
  • A CPU profile looks unremarkable, or shows time in spin loops inside lock implementations.

Two impostors produce similar curves. False sharing slows threads that touch different variables on the same cache line, with no lock involved, and shows up as high cycles per instruction rather than waiting. A saturated downstream, such as a connection pool or a database, also flattens throughput, but threads then wait on sockets or pool semaphores. Check where threads wait before assuming it is your lock.

The arithmetic of a busy lock

A lock is a single-server queue, and queueing theory tells you when it will hurt. If threads acquire a lock at rate λ per second and hold it for S seconds on average, its utilisation is ρ = λ × S. A lock acquired 40,000 times per second for 20 microseconds is busy 80 percent of the time.

Waiting time grows non-linearly with utilisation. For a rough model with random arrivals, the mean wait before acquiring is about ρ / (1 − ρ) × S. At 50 percent that is 20 microseconds of wait per acquisition; at 80 percent, 80 microseconds; at 90 percent, 180; at 95 percent, 380. That curve explains why contention seems to appear suddenly: a 15 percent increase in traffic can quadruple waiting. It also sets a hard ceiling: a lock held for 20 microseconds cannot serve more than 50,000 acquisitions per second however many cores you buy, which is the serial fraction in Amdahl's law made concrete.

So every diagnosis ends with two numbers, acquisition rate and hold time, and every fix reduces one of them or splits the lock so each part sees a fraction of the rate.

A triage workflow

The workflow below keeps you from fixing the wrong thing. Each step has a cheap check before you reach for heavier tools.

Symptomflat throughput, idle CPUIs it waiting?off-CPU, cs/s, futexNot a lockpool, I/O, false sharingWhich lock?contention profilerWhy?rate x hold timeLong holdshrink, move I/O outHigh ratestripe, batch, per-threadRead-mostlyRW lock, copy-on-writeToo many threadsshrink poolVerifyre-measure same loadnoyesConfirm waiting first, then name the lock, then measure rate and hold time; the cause picks the fix.
Lock contention triage. The two quantities in the middle box decide which branch of fixes applies.

Off-CPU analysis answers the first question across languages: it records where threads block, not where they run. The production profiling guide covers running such profilers continuously; for a one-off investigation, a runtime-specific contention profiler is usually faster.

Tools by runtime

RuntimeToolWhat it reports
Linux, any languageOff-CPU profiling, for example bcc offcputimeStacks where threads sleep, weighted by time asleep, including user-space lock waits
Linux kernelperf lock contention -a -b -- sleep 10Waits on kernel locks such as mutexes, rwsems and spinlocks, for example mmap or filesystem locks
JVMJDK Flight Recorder events jdk.JavaMonitorEnter and jdk.ThreadParkBlocked monitor entry and parked waits with durations and stacks, above a configured threshold
JVMasync-profiler in lock modeFlame graph of contended locks weighted by wait time
JVMThread dumps (jstack or jcmd Thread.print)Which threads are BLOCKED and which thread owns each monitor
GoMutex and block profiles in runtime/pprofContended sync.Mutex time and blocking on channels and locks
C and C++Off-CPU flame graphs, or timing wrappers around your locksWhere threads sleep, and per-lock wait and hold time

Two details save hours. JFR only records lock events longer than a threshold set in the recording's settings file, so short but frequent waits can be invisible until you lower it. Go's mutex profile attributes contention to the stack that releases the lock, the end of the critical section that made others wait, which is usually where the fix belongs, while the block profile shows the waiters.

Worked example: a synchronized cache in Java

A Java service keeps a cache of pricing rules in a synchronized map. Under a load test at 2,000 requests per second throughput stops rising at 16 threads, CPU sits at 35 percent of 32 cores, and p99 latency grows from 8 to 60 milliseconds. Record and inspect lock events:

jcmd <pid> JFR.start name=lock settings=profile duration=60s filename=/tmp/lock.jfr
jfr print --events jdk.JavaMonitorEnter /tmp/lock.jfr | head -50
jfr summary /tmp/lock.jfr

The events point at one monitor, the cache, held inside PricingCache.get. Each request does 20 lookups, so the lock sees about 40,000 acquisitions per second, and the stack shows why the hold time is high: on a miss, the code loads the rule from a database while holding the lock.

// Before: every lookup serialises, and a miss holds the lock across I/O.
public synchronized Rule get(String key) {
    Rule r = map.get(key);
    if (r == null) {
        r = db.loadRule(key);          // milliseconds, under the lock
        map.put(key, r);
    }
    return r;
}

// After: lock-free reads, per-key loading, no I/O under a shared lock.
private final ConcurrentHashMap<String, CompletableFuture<Rule>> map = new ConcurrentHashMap<>();

public Rule get(String key) {
    return map.computeIfAbsent(key,
            k -> CompletableFuture.supplyAsync(() -> db.loadRule(k), loader))
        .join();
}
// Also drop failed futures, or one failed load is cached for ever:
// f.whenComplete((r, e) -> { if (e != null) map.remove(key, f); });

Storing a future rather than the value matters: computeIfAbsent holds a bin-level lock while its function runs, so a slow function inside it would recreate the problem in miniature. With the future, the map operation is short, and concurrent requests for the same missing key share one load. After the change, the same test reaches 6,000 requests per second with CPU at 80 percent, and JFR shows no long monitor waits. Re-measuring under the same load is part of the fix; without it you have only moved the contention.

Go: mutex and block profiles

In Go, enable the profiles explicitly, because both are off by default, then read them with pprof:

import (
    "net/http"
    _ "net/http/pprof"
    "runtime"
)

func main() {
    runtime.SetMutexProfileFraction(5)   // sample about 1 in 5 contention events
    runtime.SetBlockProfileRate(10000)   // sample blocking events of about 10us and longer
    go http.ListenAndServe("localhost:6060", nil)
    // ...
}

// go tool pprof -top http://localhost:6060/debug/pprof/mutex
// go tool pprof -http=:8081 http://localhost:6060/debug/pprof/block

Remember the attribution rule when you read the output: the top entries in the mutex profile are the Unlock sites, the code that held the lock too long or too often.

Measuring a lock yourself

Where no profiler fits, measure the lock directly. A wrapper that records wait and hold time per named lock gives you the two numbers the diagnosis needs, and is cheap enough for a canary. A Python version, easy to port:

import threading, time
from collections import defaultdict

STATS = defaultdict(lambda: {"n": 0, "wait": 0.0, "hold": 0.0, "max_wait": 0.0})

class TimedLock:
    def __init__(self, name):
        self.name, self._lock = name, threading.Lock()

    def __enter__(self):
        t0 = time.perf_counter()
        self._lock.acquire()
        self._t1 = time.perf_counter()
        s = STATS[self.name]
        s["n"] += 1                          # safe: updated while holding the lock
        s["wait"] += self._t1 - t0
        s["max_wait"] = max(s["max_wait"], self._t1 - t0)
        return self

    def __exit__(self, *exc):
        STATS[self.name]["hold"] += time.perf_counter() - self._t1
        self._lock.release()

def report(seconds):
    for name, s in STATS.items():
        rate, hold = s["n"] / seconds, s["hold"] / max(s["n"], 1)
        print(f"{name}: {rate:.0f}/s hold={hold*1e6:.1f}us "
              f"util={rate*hold:.0%} max_wait={s['max_wait']*1e3:.1f}ms")

The utilisation column is ρ from the queueing model. Anything above about 50 percent deserves attention, and anything near 100 percent is the bottleneck.

Measurement traps

Several effects make contention look smaller, larger or different from what it is. Knowing them keeps the measurements honest.

  • Adaptive spinning hides waits. Many mutexes, including the JVM's monitors, glibc's adaptive mutex type and Go's sync.Mutex, spin briefly before sleeping. Short waits then appear as CPU time inside the lock implementation rather than as blocked time, so a CPU profile with a hot spot in lock code is contention too.
  • The lock word itself is shared data. Even when waits are short, every acquisition by a different core moves the cache line holding the lock between cores. At very high acquisition rates this transfer cost dominates, so the lock is slow while barely ever being held by two threads at once. Striping helps here, and so does keeping work on one thread.
  • Convoys. When a holder is descheduled or stalls on a page fault or a garbage collection pause, every waiter piles up behind it, and the queue can persist long after the stall ends because each wake-up costs microseconds. A single long pause can show up as minutes of elevated latency.
  • Priority and scheduling effects. A low-priority or CPU-throttled thread holding a lock delays high-priority waiters; container CPU quotas cause the same pattern when a holder is throttled mid-section.
  • Unrealistic load. A test that spreads requests evenly over keys hides the hot key a real workload has, and a test with too few client threads never builds a queue. Reproduce the production distribution of keys and concurrency.
  • Averages. Mean wait can be small while a few requests wait for tens of milliseconds. Record the maximum and a high percentile alongside the mean.

Choosing a fix

CauseFixTrade-off
I/O or allocation under the lockDo it before or after the critical section; publish results atomicallyMore careful code; possible duplicate work
Long computation under the lockCopy inputs out, compute unlocked, swap results inMust handle concurrent swaps
High rate on one lockLock striping or sharding by key; per-thread counters merged laterCross-shard operations get harder
Read-mostly dataReadWriteLock, or immutable snapshots replaced by copy-on-writeRW locks cost more per acquisition and can starve writers; see read-write locks
Many tiny updatesBatch them, or use atomic adders and lock-free structuresHarder to reason about; see lock-free structures
Too many threads for the workShrink the pool towards core countLess overlap for blocking I/O

Avoid two tempting non-fixes. Switching to a spinlock reduces context switches but burns the CPU that waiting used to free, and makes things worse under oversubscription. Making a lock fair improves tail latency for waiters but usually lowers throughput, because handing the lock to a sleeping thread costs a wake-up each time.

What to do next

  1. Confirm the signature: flat throughput, CPU below core count, rising context switches and futex calls.
  2. Take three thread dumps or an off-CPU profile under load and find the lock most threads wait on.
  3. Run the contention profiler for your runtime, lowering JFR thresholds or enabling Go's mutex profile as needed.
  4. Measure acquisition rate and hold time for the suspect lock and compute utilisation.
  5. Move I/O and slow work out of critical sections first; then stripe, batch or switch to read-optimised structures.
  6. Re-run the identical load test and compare throughput, p99 and lock utilisation before and after.
  7. Add a lock wait metric or canary profile so the next regression is caught before users notice.
Key takeaway: Lock contention shows up as flat throughput with idle CPU, rising context switches and growing tail latency. Treat each lock as a queue: utilisation is acquisition rate times hold time, and waiting explodes as it nears 100 percent. Confirm that threads are waiting, name the lock with off-CPU profiling, JFR, async-profiler or Go's mutex profile, then measure rate and hold time. Move I/O out of critical sections first, then stripe, batch or use read-optimised structures, and always re-measure under the same load.