When interviewers ask about distributed locks they want to see three things: you understand the problem space, you can compare concrete solutions, and you can tell a story that shows you used the right tool in the right context.
Why Do We Need Distributed Locks?
A lock is a way to enforce mutual exclusion – only one thread or process can hold the resource at a time. In a single‑machine program a simple mutex works because all code shares the same memory. In a cluster, however, multiple nodes may try to update the same data (e.g., a row in a database, a file in a shared store, or a counter in a microservice). Without coordination you get race conditions, lost updates, or corrupted state.
Typical scenarios that need a distributed lock include:
- Generating sequential IDs across services.
- Performing a one‑off migration or cleanup task.
- Controlling access to a limited external resource (e.g., a payment gateway quota).
Core Properties of a Correct Distributed Lock
| Property | What It Means |
|---|---|
| Safety (mutual exclusion) | No two nodes can hold the lock at the same time. |
| Liveness (progress) | If a node requests the lock and the system is healthy, it will eventually acquire it. |
| Fault tolerance | The lock must survive node crashes, network partitions, and clock drift. |
| Fairness (optional) | Requests are served in roughly the order they arrived. |
If any of these break, you risk data loss or service outages. Interviewers often probe which properties a candidate prioritizes and why.
Common Implementations
Zookeeper (or any leader‑election service)
- Uses ephemeral sequential znodes to create a queue. The node with the lowest sequence number holds the lock.
- Guarantees safety because Zookeeper itself is strongly consistent.
- Liveness depends on session timeouts; a slow network can cause a lock holder to be evicted prematurely.
- Good for scenarios where you already use Zookeeper for configuration or service discovery.
Redis RedLock (multi‑master algorithm)
- Works by acquiring the same lock key in a majority of independent Redis instances.
- Relies on loosely synchronized clocks; each lock has a short TTL.
- Provides safety only when a majority of nodes are up and clocks are reasonably aligned.
- Often chosen for low‑latency, high‑throughput workloads where a full consensus service would be overkill.
etcd (Raft‑based KV store)
- Stores a lock as a key with a lease. The lease expires if the holder crashes.
- Strong consistency from Raft gives solid safety guarantees.
- Slightly heavier than Redis but lighter than Zookeeper for many cloud‑native apps.
Database‑level locks (SELECT … FOR UPDATE, advisory locks)
- Simple to add if you already have a relational DB.
- Limited to the DB’s transaction scope; not suitable for cross‑service coordination.
- Useful for short‑lived critical sections that touch the same tables.
Typical Interview Flow
1. Basic Question
Q: What is a distributed lock and when would you use one? A (45‑90 s): "A distributed lock is a coordination primitive that guarantees only one node in a cluster can hold a resource at a time. We use it when multiple services need to modify the same piece of state – for example, generating a sequential order number across microservices. Without a lock, two nodes could read the same value, increment it, and write back the same result, causing duplicates." Follow‑up: Can you give an example from your own work?
2. Design Trade‑offs
Q: How does Redis RedLock differ from Zookeeper’s lock implementation? A: "RedLock spreads the lock across several independent Redis instances and relies on a short TTL. It’s fast and works well when you need millisecond latency, but safety hinges on a majority of nodes staying alive and clocks staying roughly in sync. Zookeeper, on the other hand, uses a strongly consistent quorum; it’s slower but gives stronger guarantees because the lock state is stored in a single ordered log. The choice usually comes down to whether you value raw speed or strict consistency." Follow‑up: What failure mode worries you most with RedLock?
3. Failure Scenarios
Q: What happens if a lock holder crashes before releasing the lock? A: "If the lock is backed by a system that can detect session loss (Zookeeper’s ephemeral nodes, etcd leases), the lock is automatically released when the session expires. With a pure Redis lock you must rely on the TTL; if the TTL is too long you risk a deadlock, if it’s too short you risk a premature release while the holder is still working. That’s why many implementations renew the TTL periodically while the lock is held." Follow‑up: How would you test that your renewal logic works under load?
4. Correctness Guarantees
Q: Do distributed locks provide linearizability? A: "Not necessarily. Some implementations, like Zookeeper, give you linearizable semantics because every operation goes through a consensus log. Others, like RedLock, provide only eventual consistency – two nodes might briefly think they both hold the lock if clocks drift. When you need strict ordering you should pick a consensus‑based store; otherwise you can accept the weaker guarantees for performance." Follow‑up: If your system only needs eventual consistency, how would you mitigate the risk of two nodes acting simultaneously?
5. Real‑World Story
Q: Tell me about a time you used a distributed lock and what you learned. A (spoken template): "In my last project we had a batch job that generated daily invoice numbers across three microservices. The initial design used a simple database auto‑increment, but we hit a duplication bug when two services started at the same time after a deploy. I introduced an etcd‑based lock around the number generation code. Each service would acquire the lock, read the current counter, increment it, and write it back, then release the lock. The lock lease automatically cleared if a service crashed, so we never got stuck. After the change the duplicate rate dropped to zero, and the latency impact was under 5 ms, which was acceptable for our SLA. The biggest lesson was to keep the lock TTL short and to renew it only while the critical section was truly active – otherwise we introduced unnecessary contention." Follow‑up: What would you do differently if the job needed to scale to hundreds of nodes?
Sample Answer Templates
Below are concise, spoken‑style templates you can adapt. Keep them under a minute, focus on clarity, and tie back to your own resume when possible.
- Basic definition: "A distributed lock ensures exclusive access to a shared resource across multiple machines. We need it when a piece of state can be changed by more than one service at a time, like generating sequential IDs."
- Implementation comparison: "RedLock spreads the lock across several Redis nodes and uses a short TTL – it’s fast but depends on clock sync and a majority being up. Zookeeper uses a quorum of servers and stores the lock as an ordered node, which is slower but gives stronger safety guarantees."
- Failure handling: "If the holder crashes, a system with session‑aware primitives (Zookeeper, etcd) will automatically remove the lock. With plain Redis you rely on the TTL, so you must renew it frequently while you hold the lock and set the TTL short enough to avoid deadlocks."
- Real story: "In my recent role I wrapped our invoice‑number generator with an etcd lock. The lock lease auto‑released on crash, eliminating duplicate numbers and keeping latency under 5 ms."
How to practice this
- Write a short script that implements a lock with etcd or Redis, then deliberately crash the process and observe how the lock is released.
- Record yourself answering the basic definition and implementation‑comparison questions in 45‑90 seconds. Play it back and trim any filler.
- Use Call Assistant to simulate a mock interview: let it detect the question, feed it your answer, and have it suggest a plausible follow‑up so you can rehearse the next turn.
FAQ
- What is the main advantage of using Zookeeper over Redis for locks? Zookeeper provides strong consistency via a consensus protocol, which guarantees that only one client can hold the lock at any time, even under network partitions.
- Can a distributed lock be completely lock‑free? No. The purpose of a lock is to serialize access; lock‑free algorithms replace locks with atomic primitives but still need a coordination mechanism for cross‑node exclusivity.
- How do you avoid the “thundering herd” problem when many nodes wait for a lock? Use exponential back‑off with jitter, and consider a queue‑based lock (like Zookeeper’s sequential nodes) that wakes only the next waiter.
- Is a TTL always required for a distributed lock? Not for systems that can detect session loss (e.g., Zookeeper, etcd). For plain key‑value stores like Redis you need a TTL to prevent deadlocks if the holder fails.
Frequently asked questions
What is the main advantage of using Zookeeper over Redis for locks?
Zookeeper provides strong consistency via a consensus protocol, which guarantees that only one client can hold the lock at any time, even under network partitions.
Can a distributed lock be completely lock‑free?
No. The purpose of a lock is to serialize access; lock‑free algorithms replace locks with atomic primitives but still need a coordination mechanism for cross‑node exclusivity.
How do you avoid the “thundering herd” problem when many nodes wait for a lock?
Use exponential back‑off with jitter, and consider a queue‑based lock (like Zookeeper’s sequential nodes) that wakes only the next waiter.
Is a TTL always required for a distributed lock?
Not for systems that can detect session loss (e.g., Zookeeper, etcd). For plain key‑value stores like Redis you need a TTL to prevent deadlocks if the holder fails.
#distributed locks#interview prep#systems design#consensus#failure handling#concept questions