Paxos is one of the core building blocks for reliable distributed systems. When an interviewer asks you to explain it, they’re looking for a clear mental model, an awareness of its limits, and the ability to map it to real projects you’ve worked on. Below is a compact way to cover everything they expect.

One‑sentence definition

Paxos is a consensus algorithm that lets a set of distributed nodes agree on a single value, even if some nodes crash or messages are delayed.

How the mechanism works

At its heart Paxos follows a three‑phase process that repeats until a value is chosen.

1. Prepare (proposal)

  • A proposer picks a unique proposal number and sends a Prepare message to a majority of acceptors.
  • Acceptors respond with the highest-numbered proposal they have already accepted (if any) and promise not to accept lower numbers.

2. Accept (vote)

  • The proposer, after receiving a majority of promises, sends an Accept request with the proposed value (or the highest‑numbered value it learned).
  • Acceptors that have not promised a higher number record the value and send an acknowledgment.

3. Learn (commit)

  • Once the proposer collects acknowledgments from a majority, it broadcasts a Learn message so all nodes can apply the chosen value.

The algorithm guarantees safety (no two different values can be chosen) because any two majorities intersect. Liveness (eventual decision) depends on having a stable leader and reliable communication.

Trade‑offs to discuss

AspectBenefitCost
ConsistencyGuarantees a single agreed‑upon value across the cluster.Requires a majority quorum, so a minority of nodes cannot make progress.
Fault toleranceCan survive failures of up to half the nodes (minus one).Latency can increase when network partitions force retries.
Simplicity of core ideaConceptually clean: prepare‑accept‑learn.Real‑world implementations add leader election, log replication, and optimizations that increase complexity.
ScalabilityWorks well for small to medium clusters (e.g., 3‑7 nodes).Larger clusters suffer from higher quorum communication cost.

Mention that many production systems (e.g., Google’s Chubby, etcd) use Paxos‑style protocols, but they often layer additional features like snapshots and membership changes.

Concrete example you can tell

Imagine a microservice that stores a configuration flag used by several instances. The flag must be the same for all instances, even if one instance crashes while updating it.

  1. Node A wants to change the flag from false to true. It generates proposal number 5 and sends Prepare(5) to nodes B, C, and D.
  2. Nodes B, C, D reply with promises, noting they have not accepted any proposal yet.
  3. Node A sends Accept(5, true) to the same three nodes.
  4. B, C, D store the value true and acknowledge.
  5. After receiving a majority (3 out of 4), Node A broadcasts Learn(true), and all nodes apply the new flag.

If Node C crashes after step 2, the remaining nodes (A, B, D) still form a majority, so the change succeeds. If a network glitch delays messages, the proposer may retry with a higher number, preserving safety.

Typical interview questions

  1. Why does Paxos need a majority quorum? – To guarantee intersecting sets of acceptors, ensuring only one value can be chosen.
  2. What happens if two proposers issue overlapping numbers? – The one with the higher number wins; lower‑numbered proposers will be rejected and must retry.
  3. How does leader election improve liveness? – A stable leader can skip the Prepare phase for subsequent proposals, reducing round‑trips.
  4. What are the differences between Basic Paxos and Multi‑Paxos? – Multi‑Paxos batches many log entries under a single leader, amortizing the Prepare cost.
  5. Can Paxos tolerate network partitions? – It can make progress only in the partition that holds a majority; the minority must wait.
  6. How would you test a Paxos implementation? – Use fault‑injection to drop messages, kill nodes, and verify that the chosen value never diverges across replicas.

60‑second spoken answer (ready for Call Assistant practice)

"Paxos is a consensus protocol that lets a group of nodes agree on a single value despite crashes or delayed messages. It works in three steps: a proposer sends a Prepare with a unique number to a majority; the acceptors promise not to accept lower numbers and return any previously accepted value. The proposer then sends an Accept with the chosen value, and once a majority acknowledges, the value is broadcast as learned. The algorithm guarantees safety because any two majorities intersect, so only one value can be chosen. It tolerates up to half the nodes failing, but it incurs latency because each decision needs a quorum round‑trip, and liveness depends on a stable leader. In practice, systems like etcd use a Multi‑Paxos variant to batch log entries and reduce the Prepare overhead."

How to practice this

  1. Write the three‑phase flow on a whiteboard and rehearse it aloud until you can cover it in under a minute.
  2. Create a tiny simulation (e.g., a Python script) that runs Prepare/Accept/Learn among three nodes, then inject failures to see safety hold.
  3. Use Call Assistant to record yourself answering the 60‑second version, then listen back for pacing and filler words, adjusting until the answer feels natural.

FAQ

  • What is the main difference between Paxos and Raft? Paxos is a family of algorithms focused on the consensus core; Raft builds on the same safety guarantees but adds a more explicit leader election and log replication model that many find easier to reason about.
  • Is Paxos still used in 2026? Yes, many distributed lock services and configuration stores continue to rely on Paxos‑style protocols, often wrapped in higher‑level APIs.
  • Can Paxos handle dynamic membership changes? Basic Paxos assumes a static set of acceptors, but extensions like Flexible Paxos or reconfiguration protocols allow nodes to join or leave while preserving safety.
  • Why do some engineers prefer Multi‑Paxos over Basic Paxos? Multi‑Paxos amortizes the costly Prepare phase across many log entries, reducing latency for steady‑state operations.

Frequently asked questions

What is the main difference between Paxos and Raft?

Paxos is a consensus core that can be built in many ways; Raft adds a clear leader election and log replication design that many find easier to understand and implement.

Is Paxos still used in 2026?

Yes, distributed lock services, configuration stores, and some database replication layers still rely on Paxos‑style protocols, often behind higher‑level abstractions.

Can Paxos handle dynamic membership changes?

The basic algorithm assumes a fixed set of acceptors, but extensions such as Flexible Paxos or reconfiguration protocols allow nodes to be added or removed while preserving safety.

Why choose Multi‑Paxos over Basic Paxos?

Multi‑Paxos reuses a stable leader to avoid repeating the Prepare phase for every command, cutting latency for workloads that issue many sequential updates.

#concept#Paxos#distributed-systems#consensus#interview