Raft is a consensus algorithm that makes a group of servers appear as a single, reliable state machine. In plain words, it lets a distributed system agree on the same sequence of commands even when some machines fail or messages are delayed. The design goal is readability – you should be able to walk through the whole protocol on a whiteboard in an hour.
The Core Mechanism: Leader, Terms, and Log Replication
Leader Election
When the cluster starts, or when the current leader crashes, the servers hold an election. Each server maintains a term number that increases monotonically. A server that times out without hearing from a leader becomes a candidate, increments its term, and asks other servers for votes. A candidate wins if it receives votes from a majority of the cluster. Once elected, it assumes the role of leader for that term.
Log Replication
The leader receives client commands, appends them to its own log, and then sends AppendEntries RPCs to followers. Each follower stores the entries in the same order. An entry is considered committed when the leader knows that a majority of servers have stored it. At that point the leader applies the entry to its state machine and informs the client.
Safety Guarantees
Raft guarantees that once a log entry is committed, no other entry with a different command can appear in the same position in any future log. This is achieved by two rules:
- Leader Completeness – a leader must have all entries that were committed in previous terms.
- Log Matching – if two logs contain an entry at the same index with the same term, the entries that follow are identical.
Trade‑offs Compared to Other Consensus Protocols
| Aspect | Raft | Paxos (classic) |
|---|---|---|
| Understandability | Designed for teaching; whole protocol fits on a whiteboard. | Conceptually simple but many edge cases make it hard to implement correctly. |
| Leader Role | Explicit leader handles all client traffic, simplifying client interaction. | No distinguished leader in basic Paxos; clients may contact any proposer. |
| Performance | Stable leader gives low latency; leader change can cause a brief pause. | Similar latency when a stable proposer exists; more frequent quorum changes can increase latency. |
| Failure Handling | Requires a majority to elect a new leader; during a split‑brain only the majority proceeds. | Also needs a majority, but without a clear leader the system can appear more chaotic. |
| Implementation Complexity | Moderate – clear state machine, term handling, and snapshotting. | Higher – need to manage multiple concurrent proposals and potentially more complex recovery logic. |
Overall, Raft trades a little extra latency during leader changes for a protocol that is easier to reason about and to debug. That simplicity is why many modern databases and key‑value stores (e.g., etcd, Consul) choose Raft.
Concrete Example: A Three‑Node Cluster
Imagine three servers: A, B, and C. All start in term 0 with empty logs.
- Election – A times out first, becomes candidate, increments term to 1, and asks B and C for votes. Both grant it, so A becomes leader.
- Client Request – A receives
SET x=5. It appends the command as log entry 1, term 1, then sends AppendEntries to B and C. - Replication – B and C store entry 1 and reply with success. A now knows a majority (A, B, C) have the entry, so it marks it committed and applies it.
- Failure – A crashes. B times out, becomes candidate for term 2, gets C’s vote, becomes leader. The new leader continues from the last committed entry, ensuring no duplication.
This tiny scenario illustrates the three pillars: election, replication, and commitment.
Typical Interview Questions About Raft
- Explain Raft in one sentence. Answer: "Raft is a leader‑based consensus algorithm that ensures a replicated log stays consistent across a cluster by having a leader replicate entries to a majority of followers."
- How does leader election work? Discuss term numbers, election timeouts, voting, and the requirement for a majority.
- What happens if the leader crashes during log replication? Explain that followers will detect a missing heartbeat, trigger a new election, and the new leader will only commit entries that a majority already stored.
- Why does Raft require a majority to commit an entry? Because a majority guarantees that any two committed entries overlap, preserving safety across terms.
- How does Raft handle log inconsistencies? Mention the AppendEntries consistency check (prevLogIndex/prevLogTerm) and how the leader overwrites conflicting entries.
- What are the main drawbacks of Raft? Talk about the brief pause during leader changes, the need for a stable network for low latency, and the fact that a single leader can become a bottleneck in write‑heavy workloads.
A 60‑Second Spoken Answer
"Raft is a consensus protocol that makes a group of servers behave like one reliable state machine. It does this by electing a single leader for each term; the leader receives client commands, appends them to its own log, and replicates the entries to the followers via AppendEntries RPCs. An entry is considered committed once a majority have stored it, at which point the leader applies it to its state machine and replies to the client. If the leader fails, the remaining servers start a new election, ensuring that only one leader exists at a time. The algorithm trades a brief pause during leader changes for a design that is easy to understand and debug, which is why many modern services choose Raft for replication."
Practicing this answer aloud, perhaps with Call Assistant listening and giving you instant feedback, can help you keep the timing and flow natural.
How to Practice This
- Write the answer on a whiteboard – sketch the three‑node example and label terms, votes, and log entries.
- Record yourself for 60 seconds – use a voice recorder or a tool like Call Assistant to capture the answer and replay it, trimming any filler.
- Run mock interview questions – have a friend ask the typical questions above and answer them using the same concise style.
FAQ
- What is the difference between Raft and Paxos? Raft explicitly defines a leader and separates the protocol into clear phases (election, replication, safety), making it easier to implement and explain. Classic Paxos focuses on agreement without a designated leader, which can be harder to reason about in practice.
- Can Raft handle network partitions? Yes. If a partition contains a majority, it can elect a leader and continue processing. The minority partition cannot make progress until it rejoins the majority.
- How does Raft ensure that committed entries are never lost? Because an entry is only committed after a majority acknowledges it, any later leader will have that entry in its log, guaranteeing durability.
- Is Raft suitable for high‑throughput write workloads? It works well for moderate write loads, but if the leader becomes a bottleneck, systems often shard data or use multiple Raft groups to scale horizontally.
Frequently asked questions
What is the one‑sentence definition of Raft?
Raft is a leader‑based consensus algorithm that keeps a replicated log consistent across a cluster by having a leader replicate entries to a majority of followers.
How does Raft handle a leader crash during replication?
Followers detect the missing heartbeat, trigger a new election, and the new leader only commits entries that a majority already stored, preserving safety.
Why do many modern services choose Raft over Paxos?
Raft’s explicit leader and clearly separated phases make it easier to understand, implement, and debug, which reduces operational risk compared to classic Paxos.
What are the main trade‑offs of using Raft?
Raft offers simplicity and predictable performance but incurs a brief pause during leader changes and can become a write bottleneck if a single leader handles all traffic.
#concept#the Raft consensus algorithm#distributed systems#interview prep#consensus