A deadlock is the quietest failure in concurrent software. Nothing crashes, nothing logs and CPU usage drops. Requests that touch the stuck code path time out one by one until a health check or an on-call engineer notices. The textbook definition is easy: a set of threads, each waiting for something that another member of the set holds. Real deadlocks are harder to spot because they rarely look like the textbook diagram. They hide in callbacks, thread pools, class initialisers and async code that blocks.
This article covers the practical side. It catalogues the shapes deadlock takes in production code, walks one bug from symptom to fix, and shows how to read the evidence in Java, Python and Go. It also covers the deadlocks that the built-in detectors cannot see and the test-time tools that catch lock-order bugs before they ship. The theory of prevention, avoidance and detection (the Coffman conditions, wait-for graphs, the Banker's algorithm and recovery strategies) is covered in deadlock detection and prevention architecture, and this page assumes it only lightly.
One model for every deadlock
You need one model to reason about every case below. Draw a node for each thread and each resource. Draw an edge from a resource to the thread that owns it, and an edge from a thread to the resource it is blocked on. A deadlock exists exactly when this graph has a cycle and every resource in the cycle is held exclusively. The word resource is deliberately broad. A mutex counts, and so does a worker slot in a thread pool, a database row lock, a connection in a pool, a bounded queue's free capacity, or the single thread of an event loop.
Thinking in resources rather than locks is the most useful habit in this article. Most deadlocks that survive code review are cycles in which at least one edge is not a lock at all.
Seven shapes deadlock takes in real code
Seven shapes account for nearly every deadlock seen in application code:
- Lock-order inversion. Two code paths take the same two locks in opposite orders. Transfers between accounts, moving an item between two caches and updating both ends of a graph edge are typical. It only fires when the two paths interleave, so it can survive years of testing.
- Calling alien code while holding a lock. A method notifies listeners, invokes a callback or logs through a custom appender while still inside its critical section. The callee takes another lock, and somewhere else that lock is held by a thread waiting for yours. The lock order is fine inside each class; the inversion spans the boundary between them.
- Self-deadlock on a non-reentrant lock. Python's
threading.Lock, Go'ssync.Mutexand POSIX default mutexes are not reentrant. A method that holds the lock and calls another public method of the same object, which locks again, waits for itself forever. Java'ssynchronizedandReentrantLockhide this class of bug, butStampedLockand semaphores do not. - Thread-pool starvation deadlock. A task running in a bounded pool submits a subtask to the same pool and blocks on its future. When every worker is doing the same thing, no worker remains to run the subtasks. No lock is involved, so lock-based detectors report nothing.
- Resource-pool cycles. A request holds one database connection and asks for a second, for example by opening a nested transaction in a separate session. With a pool of 10 and 10 concurrent requests, each holding one connection, all wait forever. The same pattern occurs with semaphores and with bounded queues whose producers are also consumers.
- Class-initialisation deadlock (JVM). Class
A's static initialiser touches classBwhile another thread is initialisingB, whose initialiser touchesA. The JVM serialises class initialisation with an internal lock, so both threads wait. Thread dumps show them inRUNNABLEor waiting on class init rather than on a monitor. - Blocking on the thread that must deliver the result. Code on a single-threaded event loop, a UI thread or a scheduler with one worker calls a blocking
.get()on a future that is completed by a callback scheduled on that same thread. Sync-over-async in .NET with a synchronisation context is the well-known example, and Node, Netty and asyncio have equivalents.
Worked example: the transfer deadlock, from symptom to fix
Here is the bug in its smallest form. A payments service moves money between two accounts and locks both so the balance check and the two updates are atomic:
class Account {
final long id;
long balance;
Account(long id, long balance) { this.id = id; this.balance = balance; }
}
void transfer(Account from, Account to, long amount) {
synchronized (from) { // thread 1 locks A, thread 2 locks B
synchronized (to) { // thread 1 waits for B, thread 2 waits for A
if (from.balance < amount) throw new InsufficientFunds();
from.balance -= amount;
to.balance += amount;
}
}
}Unit tests pass, because each test runs one transfer at a time. In production, user A pays user B at the same moment that B refunds A. Thread 1 holds A's monitor and wants B's, thread 2 holds B's monitor and wants A's, and the diagram's cycle closes. The symptom is a slowly growing count of stuck request threads, and only for requests that touch those two accounts. Everything else keeps working, which is why such deadlocks often go undiagnosed for hours.
Capturing a thread dump with jcmd <pid> Thread.print (or jstack <pid>) ends with a section the JVM computes for you. Abridged, it reads:
Found one Java-level deadlock:
=============================
"transfer-2":
waiting to lock monitor 0x00007f3a1c004e08 (object 0x000000071a2b3c40, a Account),
which is held by "transfer-1"
"transfer-1":
waiting to lock monitor 0x00007f3a1c006f58 (object 0x000000071a2b3d10, a Account),
which is held by "transfer-2"The fix is to give the two locks a global order that every caller uses. The account id already provides one, so you lock the lower id first whatever the direction of the transfer:
void transfer(Account from, Account to, long amount) {
if (from.id == to.id) throw new IllegalArgumentException("same account");
Account first = from.id < to.id ? from : to;
Account second = from.id < to.id ? to : from;
synchronized (first) {
synchronized (second) {
if (from.balance < amount) throw new InsufficientFunds();
from.balance -= amount;
to.balance += amount;
}
}
}Two details matter. The self-transfer guard prevents the code from treating one account as two. And the ordering key must be unique and stable. Where no natural order exists, or the second lock is chosen dynamically, use ReentrantLock.tryLock with a timeout. Release everything on failure and retry after a randomised backoff so that two threads do not retry in lockstep. The ordered version cannot deadlock. The tryLock version cannot hang, but under contention it can livelock and burn CPU, so give it a retry budget and a metric.
Reading the evidence in each runtime
The first job during an incident is to prove that you are looking at a deadlock and not at slow I/O or lock contention. The signature is threads blocked for a long time with near-zero CPU, all on the same few locks, and a dump that does not change between captures. Take three dumps about ten seconds apart. Threads that are merely contended move on; deadlocked ones show identical stacks every time.
Java. jcmd <pid> Thread.print -l includes ownable synchronizers, so ReentrantLock owners appear as well as monitor owners. To detect deadlocks inside the process, for a health check or an alert, ask the management bean:
ThreadMXBean mx = ManagementFactory.getThreadMXBean();
long[] ids = mx.findDeadlockedThreads(); // monitors AND ownable synchronizers; null if none
if (ids != null) {
for (ThreadInfo ti : mx.getThreadInfo(ids, true, true)) {
log.error("deadlocked: {} waiting on {} owned by {}",
ti.getThreadName(), ti.getLockName(), ti.getLockOwnerName());
}
deadlockGauge.set(ids.length);
}Prefer findDeadlockedThreads to the older findMonitorDeadlockedThreads, which only sees synchronized monitors. Logging ThreadInfo.toString() is a common mistake because it truncates each stack to a handful of frames. Iterate getStackTrace() yourself if you want the full stack. The call is not free on a process with thousands of threads, so run it every 30 to 60 seconds rather than on every probe.
Python. py-spy dump --pid <pid> prints every thread's Python stack with only a brief pause. Inside the process, faulthandler.dump_traceback_later(timeout, repeat=True) works as a cheap watchdog that writes all stacks to stderr if a request has not cancelled it in time. CPython has no built-in deadlock detector, so read the cycle off the stacks of threads parked in acquire.
Go. The runtime prints fatal error: all goroutines are asleep - deadlock! only when every goroutine is blocked. A server with a live listener goroutine never trips it, so treat it as a test-time aid. In production, send SIGQUIT or fetch /debug/pprof/goroutine?debug=2. Each goroutine header shows its wait reason and how long it has been blocked, for example [sync.Mutex.Lock, 14 minutes] in recent versions or [semacquire, 14 minutes] in older ones.
Deadlocks the detectors miss
Built-in detectors only see locks they understand. These cases produce no Found one Java-level deadlock line and no Go fatal error:
- Pool starvation. Workers park in
FutureTask.getorCompletableFuture.join. Look for every thread of one pool waiting on a future and a non-empty work queue. The fix is a separate pool for subtasks, or making the parent task asynchronous (thenCompose) so it releases its worker. - Connection-pool and semaphore cycles. Threads wait in the pool's borrow method. The tell is a pool metric showing every connection checked out with zero queries in flight on the database side.
- Distributed and database deadlocks. A thread blocked on a socket read waiting for a service that waits for a lock this thread holds. Databases such as PostgreSQL and InnoDB detect their own lock cycles and abort one transaction with a deadlock error. That error is a signal to fix the ordering of your statements, not something to retry blindly forever.
- Class-initialisation deadlocks. The threads look runnable or wait on an internal init lock. The tell is two stacks each inside a different
<clinit>. - Lost wakeups. A thread waits on a condition whose signal was sent before it started waiting, or whose predicate was not rechecked in a loop. Strictly this is not a cycle, but the symptom is identical and the dump shows a lone waiter with no owner. Always wait in a
while (!predicate)loop while holding the associated lock.
Catching lock-order bugs before production
Lock-order bugs are cheap to catch before production if you look for them deliberately:
- ThreadSanitizer. Building C or C++ code with TSan (
-fsanitize=thread) records the order in which each thread acquires mutexes. It reportslock-order-inversion (potential deadlock)when two paths acquire the same pair in opposite orders, even if the test run never actually hung. Go's-raceflag uses the same runtime but reports data races only, so Go code needs lock ranks instead. Linux's lockdep does the lock-order check for kernel locks. - Lock-rank assertions. Give every lock a numeric rank and keep a thread-local stack of held ranks in debug builds. Acquiring a lock whose rank is not higher than the top of the stack throws immediately. It costs a few lines in a lock wrapper and turns a rare production hang into a deterministic test failure.
- Open calls. Make it a code-review rule that no callback, listener, logger with custom appenders or remote call runs inside a critical section. Copy the data you need under the lock, release it, then call out.
- Stress tests with jitter. Run the concurrent paths in loops with random sleeps injected between acquisitions.
- Timeouts on every wait.
tryLock(timeout),future.get(timeout), pool borrow timeouts and statement timeouts do not prevent a deadlock, but they turn a silent hang into an error you can count and alert on.
Trade-offs between remedies
Each remedy has a price. Global lock ordering is free at runtime but must be enforced by convention, and it breaks the first time someone adds a lock without knowing the order. Coarsening to one lock removes cycles by construction but serialises unrelated work, so throughput falls as core count rises. TryLock with backoff tolerates dynamic lock sets but adds retry latency and a livelock risk. Lock-free structures and single-writer designs, such as an actor or a per-shard event loop, remove locks altogether but move the problem to queue capacity, where a full queue behaves like a held lock. Read-write locks reduce contention but add their own trap: upgrading a read lock to a write lock is a self-deadlock in most implementations, as the read-write lock article explains. Pick the simplest remedy that makes the cycle impossible rather than merely unlikely. If a liveness probe restarts a pod on a detected deadlock, dump the threads to the log first, or the restart destroys the evidence.
What to do next
- Grep your codebase for nested
synchronized,lock()andwith lockblocks, list every pair of locks taken together, and write down the order each pair uses. - Pick an ordering rule (by id, by rank or by layer), apply it to every pair you found, and add a debug-build rank assertion in your lock wrapper.
- Search for calls to listeners, callbacks, loggers and remote clients inside critical sections and convert them to open calls.
- Check every bounded pool for tasks that block on futures from the same pool, and split those pools.
- Add a periodic
findDeadlockedThreadscheck (or a goroutine-dump watchdog in Go) that logs full stacks and emits a metric. - Put timeouts on every lock acquisition, future wait and pool borrow that sits on a request path.
- Run your concurrency tests under ThreadSanitizer or with injected jitter in CI, and treat a lock-order-inversion report as a failing test.
- Rehearse the incident: deliberately deadlock a staging instance and practise capturing three dumps and reading the cycle. Then continue with locks and mutexes in depth and Java ReentrantLock for the lock primitives themselves.