Most explanations of Raft stop at the happy path: a leader is elected, it copies log entries to followers, and a majority makes them committed. That picture is correct, and this site's introduction to Raft covers it well. Real bugs live elsewhere: a vote granted before it reached disk, a read served by a deposed leader, a membership change overlapping a leader change.

This article takes the implementation view: what a node persists and when, the RPC handlers as code, the commit rule for earlier-term entries, and what serious implementations add: linearizable reads, PreVote and CheckQuorum, membership changes, snapshots and testing.

Advertisement

What Raft promises, and what it does not

Raft builds a replicated state machine. Every node holds a log of commands and applies them in log order to a deterministic state machine, such as a key-value map. If every node applies the same commands in the same order, every node reaches the same state. Raft's job is only to agree on the log.

Its safety promise holds under any timing: no node ever applies a different entry at a committed index. Its liveness promise is conditional: a cluster of 2f+1 voters makes progress while a majority, f+1 of them, can talk to each other and elections are not constantly interrupted. Three voters tolerate one failure, five tolerate two. Raft does not tolerate lying nodes, and does not by itself make reads linearizable.

Raft splits the problem into leader election, log replication and safety rules that connect them. Elections and replication are covered step by step in the leader election article and the log replication article; here we focus on the rules an implementation must not get wrong.

The state a node keeps, and when it must reach disk

Three fields are persistent: currentTerm, the latest term the node has seen; votedFor, the candidate it voted for in that term; and the log itself. Everything else, including commitIndex, lastApplied and the leader's per-follower nextIndex and matchIndex, can be rebuilt after a restart.

The rule is simple to state and easy to break: a node must make persistent state durable before sending any reply that depends on it. If a node grants a vote, replies, crashes before votedFor reaches disk, restarts and then votes for a different candidate in the same term, two leaders can be elected for one term. If a follower acknowledges an append before it is fsynced and then loses power, the leader may have counted a replica that no longer exists, and a committed entry can vanish.

Advertisement

The two handlers, written out

The whole protocol runs on two RPCs, RequestVote and AppendEntries, plus InstallSnapshot for followers that fall too far behind. The sketch below is Python-shaped pseudocode that follows the rules in the Raft paper, with the persistence points marked.

# Persistent state: must be on stable storage BEFORE any RPC reply that depends on it.
#   current_term, voted_for, log[]      (log[i].term, log[i].cmd; index 0 is a sentinel)
# Volatile state: commit_index, last_applied; on leaders also next_index[], match_index[]

def on_request_vote(req):
    if req.term > current_term:
        become_follower(req.term)            # NEW term only: persist it, clear voted_for
    up_to_date = (req.last_log_term, req.last_log_index) >= (log[-1].term, len(log) - 1)
    grant = (req.term == current_term and voted_for in (None, req.candidate_id)
             and up_to_date)
    if grant:
        voted_for = req.candidate_id
        persist(current_term, voted_for)     # fsync before replying, or you can vote twice
        reset_election_timer()
    return VoteReply(current_term, grant)

def on_append_entries(req):
    if req.term < current_term:
        return AppendReply(current_term, False)
    if req.term > current_term:
        become_follower(req.term)
    role, leader_id = FOLLOWER, req.leader_id  # same term: never clear the vote
    reset_election_timer()
    if req.prev_index >= len(log) or log[req.prev_index].term != req.prev_term:
        return AppendReply(current_term, False, hint=conflict_hint(req.prev_index))
    for i, e in enumerate(req.entries, start=req.prev_index + 1):
        if i < len(log) and log[i].term != e.term:
            truncate_log_from(i)             # only ever delete on a real conflict
        if i >= len(log):
            append_to_log(e)
    fsync_log()                              # durable before the ack counts toward a quorum
    last_new = req.prev_index + len(req.entries)
    commit_index = max(commit_index, min(req.leader_commit, last_new))
    return AppendReply(current_term, True, match_index=last_new)

def leader_advance_commit():
    for n in range(len(log) - 1, commit_index, -1):
        replicas = 1 + sum(1 for f in followers if match_index[f] >= n)
        if replicas > cluster_size // 2 and log[n].term == current_term:
            commit_index = n                 # earlier-term entries commit indirectly
            break

Four details deserve attention. First, any message with a higher term makes the receiver step down and adopt the term, whatever its role. Second, the up-to-date check compares the last log term first and the last index only as a tie-break; this is what guarantees a new leader already holds every committed entry. Third, a follower truncates its log only when an existing entry actually conflicts; blindly truncating on every append breaks when an old, delayed RPC arrives after a newer one. Fourth, the conflict hint lets the leader skip back a whole term at a time instead of one entry per round trip.

Why leaders only count replicas for their own term

The last line of leader_advance_commit is a correctness rule. A leader may only mark an entry committed by counting replicas if that entry is from its current term. Entries from earlier terms become committed indirectly, when a later current-term entry commits after them.

Here is the scenario from the paper, with five servers S1 to S5. In term 2, leader S1 writes entry X at index 2 and copies it only to S2, then crashes. S5 wins term 3 with votes from S3, S4 and itself, writes a different entry Y at index 2, and crashes before replicating it. S1 comes back, wins term 4, and copies X to S3. X is now stored on S1, S2 and S3, a majority. If S1 treated that as committed and applied it, and then crashed, S5 could win term 5: its last log term is 3, higher than the term 2 on S2, S3 and S4, so they would all vote for it. S5 would then overwrite X with Y everywhere, and an applied entry would disappear.

The fix is the current-term rule. S1 in term 4 first replicates an entry of its own term, often a no-op appended right after election. Once that entry is on a majority, no server lacking it can win an election, and everything before it, including X, is safe. It is also why a new leader cannot serve linearizable reads until its no-op commits.

The write path, end to end

Inside one Raft node: the order of operations for a writeClientPUT k=vLeader: Raft coreterm, log, commitIndexWrite-ahead logappend + fsyncState machineapply in log orderSnapshot storecompact old entriesFollower Acheck prevLogIndex/TermFollower Bappend, fsync, then ackMajority ackleader + A = 2 of 3Lagging followerInstallSnapshot12 persist3 AppendEntries4 acks5 commit6 replyreply only after apply
A write in a three-node cluster: the leader appends and fsyncs locally, replicates in parallel, commits on a majority of durable acknowledgements, applies to the state machine and only then replies. Old entries are compacted into snapshots, which also serve lagging followers.

The ordering in the figure is the contract. The client gets a success reply only after the entry is committed and applied, so a reply implies durability on a majority. If the leader crashes after committing but before replying, the client sees a timeout for a write that did happen, so retries must be idempotent: attach a client id and sequence number to each command and have the state machine remember the last sequence applied per client.

Reads: the part people get wrong

Reading from the leader's local state machine is not automatically linearizable. A leader cut off by a partition does not know it has been replaced; for up to an election timeout it may still believe it leads while a new leader on the majority side accepts writes.

There are three standard answers. The simplest puts the read through the log, costing a disk write per read. ReadIndex avoids the disk write: the leader records its commit index, confirms it is still leader by collecting heartbeat acknowledgements from a majority, waits until it has applied up to that index, and then serves the read.

def linearizable_read(key):
    if role != LEADER or not committed_entry_in_current_term():
        raise NotLeaderOrNotReady()          # new leaders append a no-op first
    read_index = commit_index
    confirm_leadership_with_heartbeat_quorum()   # one round trip, no log write
    wait_until(last_applied >= read_index)
    return state_machine.get(key)

Lease reads skip even the heartbeat round. After a majority acknowledges a heartbeat, the leader assumes no other leader can be elected until the followers' election timeouts expire, and serves reads locally until then, minus a safety margin. This is correct only if clocks advance at roughly the same rate everywhere; a paused VM can break it. etcd's raft library offers both ReadIndex and a lease-based option. Follower reads extend ReadIndex: the follower gets the commit index from the leader, applies up to it and serves locally.

Liveness in practice: PreVote, CheckQuorum and timeouts

Randomised election timeouts prevent endless split votes. Two further extensions matter in real networks. PreVote adds a trial round before a real election: a node that has lost contact asks whether peers would vote for it, without incrementing its term. Peers that still hear from a healthy leader say no. Without PreVote, a node on a flaky link repeatedly bumps its term, and each time it reconnects its higher term forces the working leader to step down.

CheckQuorum is the leader-side counterpart: a leader that has not heard from a majority within an election timeout steps down on its own. That bounds how long an isolated leader keeps accepting client requests it cannot commit, and it is a prerequisite for lease reads. Leadership transfer moves leadership before a planned restart, avoiding a full election timeout.

Membership changes and snapshots

During a membership change nodes may use different configurations, and old and new majorities might not overlap. Raft offers two methods. Joint consensus moves through an intermediate configuration in which decisions need majorities of both old and new sets, then commits the new set alone; it can change several members at once. Single-server changes add or remove one voter at a time, which guarantees any old majority and new majority overlap. A bug found after publication let overlapping changes across a leader change break that; the fix is that a new leader commits an entry in its own term, the same no-op, before proposing a change.

Add new servers as learners (non-voting members) first. An empty new voter counts toward the quorum immediately and can stall writes while it catches up; promote it once its log is close to the leader's.

Logs grow forever unless compacted. Each node periodically snapshots its state machine along with the last included index and term, and deletes log entries up to that point. When the leader no longer holds the entries a follower needs, it sends InstallSnapshot instead. Stream large snapshots in chunks and take them without blocking the apply loop.

Worked example: a partition across availability zones

A five-node cluster spans three zones: two nodes in zone A (A1 is leader), two in zone B and one in zone C. Zone A is cut off. A1 and A2 can reach each other but nobody else. With CheckQuorum, A1 notices within one election timeout that it hears from only two of five voters and steps down, so clients hitting zone A get not-leader errors rather than hanging writes. On the other side, B1, B2 and C1 hold three of five votes. One times out, wins PreVote and then the election, commits a no-op and resumes service.

When zone A reconnects, A1 and A2 see the higher term, become followers, and have any uncommitted entries from their isolated period overwritten. No client saw those entries succeed. The placement lesson: a 2-2-1 split over three zones survives losing any one zone.

Throughput, batching and testing

Raft's throughput is limited by fsync latency and round trips, not CPU. Batch many client commands into one log append and one fsync, and pipeline AppendEntries so the leader sends the next batch before the previous one is acknowledged. Systems that need more throughput shard data across many Raft groups.

Test with more than unit tests. Deterministic simulation, where the network, disk and clocks are fakes controlled by a seeded scheduler, lets you replay rare interleavings exactly. Fault-injection frameworks such as Jepsen partition networks, pause processes and skew clocks against a real cluster, and linearizability checkers like Knossos or Porcupine verify the recorded client history.

Failure modes

SymptomLikely causeFix
Two leaders in one termVote or term replied before fsyncPersist currentTerm and votedFor before replying
Committed write lost after power lossDisk cache or fsync disabledReal fsync; verify with crash tests
Stale reads during partitionsLocal reads on a deposed leaderReadIndex, or leases plus CheckQuorum
Frequent elections, term keeps risingFlaky node without PreVoteEnable PreVote; fix the link
Elections whenever disk is slowElection timeout near fsync p99Raise timeout; separate WAL disk
Writes stall after adding a nodeEmpty node added as voterAdd as learner, promote when caught up
Follower never catches upLog compacted past its positionInstallSnapshot, streamed in chunks
Duplicate effects after timeoutsClient retries non-idempotent commandClient id plus sequence number dedup

Trade-offs

ChoiceOption AOption B
Cluster size3 voters: lower latency, tolerates 1 failure5 voters: tolerates 2, slower quorum
Read pathReadIndex: one round trip, no clock assumptionLease: local reads, needs bounded clock drift
Membership changeSingle-server: simplerJoint consensus: several changes at once
ScaleOne group: simple, one leader bottleneckMany groups: scales, needs routing and rebalancing
ProtocolRaft: strong leader, easier to reason aboutMulti-Paxos variants: more flexible, see the Paxos article

What to do next

  1. Find where your Raft library or database persists term, vote and log, and confirm each is fsynced before the corresponding reply.
  2. Check whether reads in your system use the log, ReadIndex or leases, and whether CheckQuorum is enabled if leases are.
  3. Enable PreVote if your library supports it, and set the election timeout from measured fsync and network p99, not defaults.
  4. Practise adding a node as a learner, promoting it and removing an old one, and time how long writes pause.
  5. Run one partition drill against a staging cluster and confirm the minority side refuses writes quickly.
  6. Add a crash-restart test and a linearizability check of client history to your test suite.
Key takeaway: Raft agrees on a log; the hard parts are around it. Persist term, vote and log before replying, count replicas only for current-term entries and commit a no-op after election, serve reads through ReadIndex or carefully bounded leases, use PreVote and CheckQuorum for stable leadership, add members as learners and change one at a time or by joint consensus, compact with snapshots, and prove all of it with fault injection rather than trust.