Most explanations of two-phase commit stop at the diagram: prepare, vote, commit. The diagram is correct and nearly useless when you have to build one, because all the difficulty sits in the gaps between the arrows. What happens if the coordinator dies after one participant commits and before the other does? Who is allowed to decide abort for a branch that has been prepared for six hours? How do you prove your implementation handles each case rather than hoping it does?

This article answers those questions in code. It builds a small coordinator over PostgreSQL prepared transactions, adds a recovery loop that cannot race with a live coordinator, and tests it by crashing at every boundary. It then looks at MySQL XA, the heuristic outcomes of the XA standard, and Kafka's transaction protocol, which is two-phase commit in disguise. The protocol's theory, presumed abort and why it blocks are covered in Two-Phase Commit Revisited; when to avoid 2PC entirely is covered in Distributed Transactions in Practice.

Advertisement

The contract, as durable writes

CoordinatorParticipant A (orders)Participant B (billing)BEGIN + workBEGIN + workC1PREPARE TRANSACTION gid-APREPARE TRANSACTION gid-Bdurable yes vote, locks heldC2Decision row: commitINSERT ... ON CONFLICTC3COMMIT PREPARED gid-AC4COMMIT PREPARED gid-BDecision row: donesafe to garbage-collectRed labels C1-C4 mark the crash points the test matrix kills the coordinator at.
A two-participant commit. Each arrow is a network call; each box is a durable write. The crash points C1-C4 sit between them.

Strip two-phase commit down to what must be on disk and it is four kinds of durable write. Each participant writes its changes and its yes vote (in PostgreSQL, PREPARE TRANSACTION). The coordinator writes one decision. Each participant writes the outcome (COMMIT PREPARED or ROLLBACK PREPARED). The coordinator finally marks the transaction done so its record can be removed.

Two rules make it correct. First, a participant that has voted yes may not decide alone: it has promised to commit if asked, so it must keep locks and wait. Second, the decision exists exactly when the decision record is durable: not when a thread chooses it, not when the first participant hears it. Everything below is engineering to make those two rules true under crashes and concurrency.

A minimal coordinator

The coordinator below uses three PostgreSQL databases: two participants and a coordinator database holding the decision log. Participants need max_prepared_transactions above zero (it is 0, disabled, by default). Connections run in autocommit mode with explicit BEGIN, because COMMIT PREPARED cannot run inside a transaction block. Global transaction ids encode coordinator, transaction and branch, so recovery can map any prepared branch back to its decision.

import uuid, psycopg
from psycopg import sql

# On the coordinator database:
# CREATE TABLE txn_decision (txid text PRIMARY KEY, outcome text NOT NULL,
#                            done boolean NOT NULL DEFAULT false);

COORD_ID = "coord-eu1"
DECIDE = ("INSERT INTO txn_decision (txid, outcome) VALUES (%s, %s) "
          "ON CONFLICT (txid) DO NOTHING")
READ = "SELECT outcome FROM txn_decision WHERE txid = %s"

def gid(txid, branch):
    return f"{COORD_ID}:{txid}:{branch}"          # PostgreSQL limit: 200 bytes

def decide(log, txid, wanted):
    # Compare-and-set: the first writer wins, everyone obeys the stored row.
    log.execute(DECIDE, (txid, wanted))
    return log.execute(READ, (txid,)).fetchone()[0]

def finish(conn, g, outcome):
    stmt = "COMMIT PREPARED {}" if outcome == "commit" else "ROLLBACK PREPARED {}"
    try:
        conn.execute(sql.SQL(stmt).format(sql.Literal(g)))
    except psycopg.errors.UndefinedObject:
        # Finished by us earlier, OR finished by hand: PostgreSQL cannot tell
        # which, so record it for reconciliation instead of hiding it.
        log_warning("branch %s missing while applying %s", g, outcome)

def transfer(log, parts, work, hook=lambda point: None):
    txid = uuid.uuid4().hex
    prepared = []
    try:
        for name, conn in parts.items():
            conn.execute("BEGIN")
            work[name](conn)
        hook("C1")
        for name, conn in parts.items():
            conn.execute(sql.SQL("PREPARE TRANSACTION {}").format(sql.Literal(gid(txid, name))))
            prepared.append(name)
        hook("C2")
        outcome = decide(log, txid, "commit")
    except psycopg.Error:                          # a participant said no or failed
        for name, conn in parts.items():
            if name not in prepared:
                conn.execute("ROLLBACK")           # unprepared work just disappears
        outcome = decide(log, txid, "abort")
    hook("C3")
    for i, (name, conn) in enumerate(parts.items()):
        if name in prepared:
            finish(conn, gid(txid, name), outcome)
        if i == 0:
            hook("C4")
    log.execute("UPDATE txn_decision SET done = true WHERE txid = %s", (txid,))
    return outcome

Three choices in this code carry most of the correctness. The decision is an INSERT ... ON CONFLICT DO NOTHING followed by a read, so if a recovery process has already written abort for this transaction, the coordinator's commit attempt loses and it obeys the abort it reads back. Finishing a branch is idempotent: a branch that no longer exists was already finished. And the coordinator only ever rolls back unprepared work directly; anything prepared goes through the decision log.

Advertisement

Recovery that cannot race the coordinator

The textbook recovery rule under presumed abort is 'no commit record means abort'. Implemented naively it has a race: a slow coordinator thread has prepared both branches and is about to write commit, while a recovery sweep sees prepared branches with no record and rolls them back. The coordinator then writes commit and tells participant B to commit a branch that A has already rolled back.

The compare-and-set decision row removes the race without locks or leases. Recovery does not roll back because a record is missing; it first tries to write abort. If it wins, nobody can ever write commit for that transaction. If it loses, it reads the commit and finishes the branches forward.

def recover(log, parts, min_age="5 minutes"):
    for name, conn in parts.items():
        rows = conn.execute(
            "SELECT gid FROM pg_prepared_xacts "
            "WHERE gid LIKE %s AND prepared < now() - %s::interval",
            (COORD_ID + ":%", min_age)).fetchall()
        for (g,) in rows:
            _, txid, _branch = g.split(":")
            outcome = decide(log, txid, "abort")   # wins only if nobody decided
            finish(conn, g, outcome)
    # A decision is done once no participant still holds a branch for it.
    mark_done_where_no_branches_remain(log, parts)

The age threshold is a performance guard, not a correctness one: even a sweep that runs on a fresh transaction is safe, it just aborts work that would otherwise have committed. Keep the decision rows until every branch is finished and the row is marked done; deleting them earlier turns a later sweep's 'abort' into a wrong answer for a transaction that actually committed.

Testing every crash point

Two-phase commit bugs live in rare interleavings, so do not wait for production to find them. The hook argument above exists for tests: raise at a named point to simulate the coordinator dying there, run recovery, and assert the invariants.

Crash pointState left behindCorrect outcomeWho finishes it
C1, before any prepareOpen transactions on both participantsAbortPostgreSQL itself, when the session ends
C2, all prepared, no decisionTwo prepared branches, no rowAbortRecovery wins the abort CAS
C3, decision loggedTwo prepared branches, commit rowCommitRecovery reads commit, finishes both
C4, one branch committedOne prepared branch, commit rowCommitRecovery finishes the second
Participant crash after preparePrepared branch survives restartAs decidedCoordinator retry or recovery
Participant crash before prepareIts work is lost, prepare failsAbortCoordinator's except path
import pytest

class Crash(Exception): pass

@pytest.mark.parametrize("point", ["C1", "C2", "C3", "C4"])
def test_crash_points(point, log, parts, work):
    def hook(p):
        if p == point:
            raise Crash(p)
    with pytest.raises(Crash):
        transfer(log, parts, work, hook)
    reopen_all(parts)                    # a crashed process loses its sessions
    recover(log, parts, min_age="0 seconds")
    a = count_rows(parts["orders"]); b = count_rows(parts["billing"])
    assert a == b, "atomicity violated"            # both applied or neither
    assert no_prepared(parts), "branch left in doubt"

The coordinator's handler catches only database errors, so the test's Crash exception escapes exactly as a dead process would: no abort is written, no rollback is sent, and recovery has to cope with whatever was left on disk. The C2 case is the one that matters most, because it is the classic blocking window: both branches are prepared and nobody has decided. Once the matrix passes in-process, repeat it with a subprocess that is killed with SIGKILL at each point, which also exercises connection loss on the participant side. Run the matrix against real databases in CI, then extend it with network partitions and paused processes; the lessons in what Jepsen taught us apply directly: the bugs are in the timing you did not test.

The same protocol in MySQL: XA

MySQL exposes the participant side through XA statements. A branch is identified by a global transaction id and an optional branch qualifier, and prepared branches survive disconnects and restarts.

XA START 'tx-7f3a', 'billing';
INSERT INTO charges (id, order_id, cents) VALUES ('ch-91', 'o-55', 1299);
XA END 'tx-7f3a', 'billing';
XA PREPARE 'tx-7f3a', 'billing';

-- later, from any session, on the coordinator's instruction
XA COMMIT 'tx-7f3a', 'billing';

-- recovery: list prepared branches (formatID, gtrid_length, bqual_length, data)
XA RECOVER CONVERT XID;
XA ROLLBACK 'tx-7f3a', 'billing';

The data column concatenates gtrid and bqual; use the length columns to split it, or CONVERT XID to get a hexadecimal form that round-trips safely. The recovery logic is identical to the PostgreSQL version: list branches that belong to this coordinator, compare-and-set a decision, finish forward or backward. In MySQL 8.0, XA RECOVER needs the XA_RECOVER_ADMIN privilege, so give the recovery job its own account.

Heuristic outcomes: when a participant breaks the promise

In theory a prepared participant waits forever. In practice an operator, staring at a prepared transaction that has held locks on the orders table for three hours while the coordinator is down, runs ROLLBACK PREPARED. The XA standard calls this a heuristic decision, and when the coordinator later asks the participant to commit, the participant reports what happened instead: heuristically committed (XA_HEURCOM), heuristically rolled back (XA_HEURRB), mixed (XA_HEURMIX) or unknown (XA_HEURHAZ). The transaction manager must then tell the participant to forget the branch, and in Java's JTA the application sees exceptions such as HeuristicMixedException.

PostgreSQL cannot distinguish a branch finished by hand from one the coordinator already finished; both are simply gone, which is why the coordinator code above logs every missing branch. A heuristic outcome is not an error your code can retry away. It is a statement that atomicity has already been violated and a human must reconcile the data. Design for it anyway: log heuristic reports loudly, include enough business identifiers in the gid or a side table to find the affected records, and make the manual resolution procedure explicit, including the rule that an operator must look up the coordinator's decision before touching a prepared branch.

Kafka transactions: 2PC with a replicated coordinator

Kafka's transactional producer is a two-phase commit that most people use without noticing. The producer registers a transactional.id; a transaction coordinator on a broker assigns it a producer id and an epoch, and bumping the epoch fences old instances of the same producer. As the producer writes to partitions it registers them with the coordinator, which records them in the internal __transaction_state topic.

On commit, the coordinator writes a prepare-commit entry to its transaction log: that replicated write is the decision. It then writes commit markers into every partition the transaction touched and finally records the transaction as complete. Consumers using isolation.level=read_committed only read up to the last stable offset, so they never see records of transactions that are still open or aborted.

Two differences from database 2PC are instructive. The participants never vote: the data is already durably appended, so 'prepare' is implicit and the only thing left to decide is visibility. And the decision log is a replicated Kafka partition, so a coordinator crash moves leadership instead of blocking. Exactly-once processing built on this is covered in exactly-once semantics. KIP-939 proposes letting a Kafka producer take part in an external 2PC as a participant that can keep a prepared transaction across restarts; check whether your broker and client versions implement it before designing around it.

Operating 2PC

  • Alert on age, not just count. On every participant, export the number and the oldest age of prepared branches (pg_prepared_xacts, XA RECOVER). A prepared PostgreSQL branch also holds back vacuum's horizon, so an old one becomes a bloat incident as well as a lock incident.
  • Run recovery continuously. A recovery loop that only runs at coordinator startup leaves branches stuck when the coordinator never restarts. Run it on a timer, on more than one host; the compare-and-set makes concurrent sweeps safe.
  • Make the decision log as durable as the data. If the coordinator database is less available than the participants, it becomes the blocking point. Replicate it with synchronous replication, or use a consensus-backed store.
  • Bound the lock time. Every millisecond between prepare and commit is lock time on hot rows. Keep the work inside prepared transactions small and avoid calling slow external services between phases.
  • Never let a script finish branches without the log. Give operators a tool that reads the decision and finishes branches; remove the temptation to run ROLLBACK PREPARED by hand.

When to use it, and when not

Two-phase commit is worth its cost when a small number of resources you control must change atomically, the participants support prepared state, and you can tolerate blocking during coordinator outages. It is the wrong tool across organisational boundaries, for long-running workflows, or where participants cannot hold locks for seconds; there sagas with compensations or an outbox with idempotent consumers give up isolation in exchange for availability. Distributed databases that run consensus-replicated commit internally give you the same atomicity with far less of this machinery exposed.

What to do next

  1. Enable prepared transactions (max_prepared_transactions in PostgreSQL) only on the databases that participate, and size it to expected concurrency.
  2. Implement the decision as a compare-and-set row and make branch finishing idempotent before writing anything else.
  3. Encode coordinator id, transaction id and branch in every gid so recovery can find the decision.
  4. Write the crash-point test matrix (C1-C4 plus participant crashes) and run it against real databases in CI.
  5. Deploy a recovery loop on a timer on at least two hosts and alert on prepared-branch age on every participant.
  6. Document the heuristic-outcome procedure and give operators a tool that finishes branches from the decision log.
  7. Revisit whether sagas, an outbox or a distributed database would remove the need for 2PC in this path.
Key takeaway: Two-phase commit is a handful of durable writes and two rules: a prepared participant cannot decide alone, and the decision exists only when its record is durable. Make the decision a compare-and-set row, finish branches idempotently, run recovery continuously, and prove the implementation by crashing it at every boundary. The same pattern explains MySQL XA, heuristic outcomes and Kafka's transaction coordinator.