A distributed lock is supposed to guarantee that only one process at a time does something: runs the nightly billing job, writes a particular file, sends a particular email. On one machine a mutex does this reliably because the operating system controls every thread. Across machines there is no such referee. Processes pause for garbage collection, networks delay packets for seconds, and clocks drift, so a process can believe it holds a lock long after the system has given it to someone else.

This article is a catalogue of the patterns people actually use, from database row locks to Redis to ZooKeeper and etcd, with code for each, the failure each one is exposed to, and a way to choose. The mechanisms underneath, leases and fencing tokens, have their own deep dives in leases and fencing tokens; here they are the vocabulary for comparing patterns.

Advertisement

First question: efficiency or correctness?

Martin Kleppmann's distinction is the most useful first question: is the lock for efficiency or for correctness? An efficiency lock stops two workers from doing the same expensive but harmless work twice, such as recomputing a cache entry. If it occasionally fails and both run, you lose some compute. A correctness lock protects an invariant: two workers both charging a card, or both appending to a ledger, is a bug that reaches customers.

The distinction matters because the cheap patterns are fine for efficiency and inadequate for correctness, and the correct patterns cost more in latency, infrastructure and code. Most incidents with distributed locks come from using an efficiency-grade lock for a correctness job.

Every pattern below is a lease: a lock that expires after a time-to-live unless renewed, because a lock without expiry is held forever by a crashed client. And every lease has the same hole, shown in the diagram: the holder can pause past the expiry without knowing it.

The hole every lease has

Why a lease alone is not mutual exclusionClient Alock, token 33GC pause / network stallwrite(33)Lock servicelease held by Alease expired -> granted to BClient Block, token 34write(34)Storage: highest token seen = 34write(33) arrives late -> REJECTEDWithout the token check, A's stale write lands after B's and silently overwrites it.The lock service cannot prevent this; only the resource that is written to can.
Client A pauses past its lease; B acquires the lock; A wakes and writes. A monotonically increasing fencing token checked by storage is what stops the stale write.

No client-side check closes this gap. Client A can check its lease immediately before writing, and the pause can happen between the check and the write, or the write can sit in a network buffer. The fix is to make the protected resource participate: each grant carries a token that increases with every grant, every write carries the token, and the resource rejects any token lower than the highest it has already accepted. A pattern that cannot produce such a token cannot be used for correctness unless the resource itself is the lock.

Advertisement

Pattern 1: let the database be the lock

If the data you are protecting lives in a relational database, the best lock is usually the database itself, because the lock and the write share one transaction and the hole above disappears. Row locks serialize work on a specific record:

BEGIN;
SELECT * FROM accounts WHERE id = 42 FOR UPDATE;   -- blocks other writers on this row
-- read, decide, write
UPDATE accounts SET balance = balance - 100 WHERE id = 42;
COMMIT;                                            -- lock released with the commit

For work that does not map to a row, such as a scheduled job, PostgreSQL's advisory locks take an arbitrary 64-bit key. pg_try_advisory_lock(key) is session-scoped: it is held until explicitly unlocked or until the connection ends, so a crashed worker releases it automatically. pg_try_advisory_xact_lock(key) is released at the end of the transaction:

def run_exclusive(conn, job_name, work):
    key = int.from_bytes(hashlib.sha256(job_name.encode()).digest()[:8], "big", signed=True)
    with conn.transaction():
        got = conn.execute("SELECT pg_try_advisory_xact_lock(%s)", (key,)).scalar()
        if not got:
            return "skipped: another worker holds it"
        work(conn)          # writes in this same transaction are protected

The caveat: a session-scoped advisory lock is only as good as the connection's liveness. A connection pooler in transaction mode (PgBouncer, for example) hands your session to other clients, which breaks session-scoped locks; use the transaction-scoped form there. And if the work writes to a different system, such as an external API, the transaction no longer covers it, and you are back to the general case.

Pattern 2: a compare-and-set lease row

A portable variant is a lease table updated with compare-and-set, which works on any store with conditional writes: SQL, DynamoDB condition expressions, Cassandra lightweight transactions, etcd. The row holds the owner, an expiry and a version that doubles as the fencing token:

-- acquire: succeeds only if free or expired; the version is the fencing token
UPDATE locks
   SET owner = :me, expires_at = now() + interval '30 seconds', version = version + 1
 WHERE name = 'billing-job'
   AND (owner IS NULL OR expires_at < now())
RETURNING version;

-- renew: only the current owner, only while it still holds the same version
UPDATE locks SET expires_at = now() + interval '30 seconds'
 WHERE name = 'billing-job' AND owner = :me AND version = :token;

-- release
UPDATE locks SET owner = NULL WHERE name = 'billing-job' AND owner = :me AND version = :token;

Because expiry is compared against the database server's clock, client clock skew does not matter. The version increases on every grant, so it is a valid fencing token for any resource willing to check it.

Pattern 3: single-instance Redis

Redis is the most common lock store because it is fast and already deployed. The correct single-instance pattern is one atomic command to acquire with a random owner value and an expiry, and a compare-and-delete script to release, so a client never deletes a lock that expired and was granted to someone else:

token = secrets.token_hex(16)
acquired = r.set("lock:billing", token, nx=True, px=30_000)   # SET key val NX PX 30000

RELEASE = """
if redis.call('get', KEYS[1]) == ARGV[1] then
    return redis.call('del', KEYS[1])
else
    return 0
end
"""
r.eval(RELEASE, 1, "lock:billing", token)

Renewal uses the same compare-then-act shape with pexpire in place of del. The weaknesses are structural. The random value is not monotonic, so it cannot be a fencing token. And with a replica, Redis replication is asynchronous: if the primary acknowledges the lock and fails before replicating it, the promoted replica has no lock and grants it again. For efficiency locks this is an acceptable, rare double run. For correctness it is not.

Redlock and the argument about it

Redlock, proposed by Redis's author, tries to remove the single point of failure: acquire the same key with the same random value on N independent Redis primaries (typically five), succeed only on a majority, and treat the lock as valid for the TTL minus the time acquisition took minus an allowance for clock drift. Kleppmann's 2016 critique, "How to do distributed locking", made two points that remain the core of the debate. First, Redlock depends on timing assumptions (bounded pauses, bounded network delay, bounded clock drift between nodes) that real systems violate, for example when a node's clock jumps or a client pauses after acquiring. Second, it produces no fencing token, so even a perfectly implemented Redlock cannot stop the stale write in the diagram. Antirez replied that the timing assumptions are reasonable in practice; the absence of a monotonic token is not disputed.

A fair reading: Redlock costs five Redis nodes and buys more availability than one instance, but not correctness. For efficiency, single-instance Redis is usually enough. For correctness, use something that produces a fencing token.

Pattern 4: consensus-backed locks

Coordination services built on consensus, ZooKeeper (ZAB) and etcd (Raft), replicate every lock operation to a majority before acknowledging it, so a leader failover cannot lose a granted lock, and every write gets a monotonically increasing number that serves as a fencing token.

In ZooKeeper the standard recipe creates an ephemeral sequential znode under the lock path. Ephemeral means it disappears when the client's session expires; sequential means ZooKeeper appends an increasing counter. The client holding the lowest number owns the lock. Every other client watches only the node immediately before its own, not the lock parent, so a release wakes exactly one waiter instead of all of them (the herd effect). The node's creation zxid is a usable fencing token.

etcd expresses the same idea with a lease and a transaction. Grant a lease with a TTL and keep it alive; create a key under the lock prefix attached to the lease; the key with the lowest create revision owns the lock; waiters watch the key just ahead of theirs. The revision is a cluster-wide monotonic counter, so it is the fencing token. The client libraries package this (Curator's InterProcessMutex for ZooKeeper, etcd's concurrency package for Go), and you should use them rather than reimplement the recipe. In outline:

lease = etcd.grant_lease(ttl=10)
keepalive_in_background(lease)                    # renew well before expiry
my_key = f"/locks/billing/{lease.id}"
txn(if_=[create_revision(my_key) == 0],           # only create, never overwrite
    then=[put(my_key, me, lease=lease)])
my_rev = get(my_key).create_revision
while True:
    waiters = get_prefix("/locks/billing/", sort_by="create_revision")
    if waiters[0].key == my_key:
        break                                     # owner; my_rev is the fencing token
    predecessor = last(w for w in waiters if w.create_revision < my_rev)
    wait_for_delete(predecessor.key)
do_work(fencing_token=my_rev)

Consensus does not abolish the pause problem: a session or lease can still expire while the holder is paused. What it gives you is a lock that is never granted twice at the same moment by the service, plus the token that lets the resource reject the paused holder. Google's Chubby paper calls these tokens sequencers for the same reason.

Fencing at the resource

The token is useless unless the resource checks it. In a database, store the highest token per protected entity and make the write conditional on it:

UPDATE invoices
   SET status = 'sent', fence = :token
 WHERE id = :invoice_id AND fence <= :token;     -- 0 rows updated = stale holder, abort

Object stores and external APIs often offer a weaker equivalent, such as conditional writes on an ETag or a generation number, and an idempotency key on payment APIs. When the resource offers neither, no lock pattern can make the operation safe on its own; design the operation to be idempotent instead.

Choosing a pattern, with a worked example

PatternFencing tokenSurvives store failoverGood for
DB row lock / advisory xact locknot needed if work is in the same transactionwith the database's own HAwork on data in that database
CAS lease rowyes (version)if the store is strongly consistentportable jobs; correctness with fencing
Redis single instancenono (async replication)efficiency: dedupe, rate limiting work
Redlocknopartially, under timing assumptionsefficiency where one Redis is not available enough
ZooKeeper / etcdyes (zxid / revision)yes (majority quorum)correctness, leader election, long-held locks

Worked example: a billing service runs three replicas, and the monthly invoice run must execute once. Invoices live in PostgreSQL. The simplest correct design is a transaction-scoped advisory lock around a run that writes only to that database. Then a requirement arrives to call a payment provider during the run. The transaction no longer covers the side effect, so the design changes: take a CAS lease row whose version is the token, write each invoice with fence <= :token, and pass the invoice ID as the payment provider's idempotency key. A replica that stalls for 40 seconds on a 30-second lease will now fail its next database write instead of double-charging, and a retried payment call is deduplicated by the provider.

Operating locks in production

  • TTL sizing. Long TTLs mean slow recovery after a crash; short ones mean spurious expiry during pauses. Renew at about a third of the TTL, and have the worker stop work if renewal fails rather than carry on hopefully.
  • Do not hold locks across slow calls without fencing. Every external call inside the critical section widens the pause window.
  • Lock ordering. Workers that take several locks must take them in one global order, or they deadlock until TTLs expire, which looks like random slowness.
  • Granularity. One global lock serializes everything; lock the entity, not the table.
  • Observability. Export acquisition latency, hold time, contention (failed tries) and fencing rejections. A non-zero fencing-rejection rate means pauses longer than your TTL are happening, which is worth knowing before it matters.
  • Stuck locks. Have a documented, audited way to break a lock by hand, and make it bump the token so the old holder is fenced out.

When not to lock at all

The best distributed lock is often none. Partition the work so each key has exactly one owner (consistent hashing, Kafka partitions consumed by one member of a consumer group). Make operations idempotent so running twice is harmless. Use the database's own constraints, such as a unique index on (customer_id, period), so the second invoice insert fails. Or elect a single leader for a whole class of work, as covered in leader election, which is the same lease problem solved once instead of per operation.

What to do next

  1. For every lock you have, write down whether it is for efficiency or correctness.
  2. Move correctness locks on database data into the database: row locks or transaction-scoped advisory locks.
  3. For correctness locks that cover other resources, switch to a pattern with a monotonic token (CAS lease row, etcd, ZooKeeper).
  4. Make the protected resource reject stale tokens with a conditional write, and alert on rejections.
  5. Audit Redis locks: atomic SET NX PX, random owner value, compare-and-delete release, and no correctness duty.
  6. Set TTLs from measured worst-case pauses, renew at a third of the TTL, and stop work when renewal fails.
  7. Add idempotency keys to every external side effect done under a lock.
  8. Look for a design that removes the lock: partitioned ownership, unique constraints or a single elected leader.
Key takeaway: A distributed lock is a lease, and any lease holder can pause past its expiry without knowing. For efficiency locks that is a tolerable double run, so single-instance Redis is fine. For correctness, either keep the lock and the write in one database transaction, or use a pattern that issues a monotonic fencing token, such as a compare-and-set lease row, etcd or ZooKeeper, and make the protected resource reject stale tokens. Redlock adds availability, not correctness. Often the better move is to remove the lock with partitioned ownership, idempotency or unique constraints.