When interviewers ask about distributed locks they want to see that you understand both the why and the how of coordinating state in a multi‑node system. A concise answer should start with a one‑sentence definition, then walk through the typical mechanism, outline the main trade‑offs, give a concrete real‑world example, and finally anticipate follow‑up questions. Below is a roadmap you can use to structure your response, plus a 60‑second spoken version you can rehearse.
What Is a Distributed Lock?
A distributed lock is a coordination primitive that guarantees exclusive access to a shared resource across a cluster of machines. In other words, it ensures that at any moment only one node can hold the lock and perform the protected operation, even though the nodes may be geographically dispersed and communicate only over a network.
Core Mechanism
Most implementations rely on a centralised or consensus‑based service that all nodes can talk to. The two most common patterns are:
- Centralised coordinator – A single service (e.g., Apache Zookeeper, Consul) stores lock state. Nodes create an ephemeral node or key representing the lock; the service grants the lock to the first creator and rejects subsequent attempts until the lock is released or the session expires.
- Consensus algorithm – Systems like etcd or Redis Redlock use a quorum of nodes to replicate lock state. A client writes a lock record to a majority of nodes with a TTL (time‑to‑live). If it can write to enough nodes, it considers the lock acquired.
Both patterns share three essential steps:
| Step | Centralised coordinator | Consensus‑based lock |
|---|---|---|
| Acquire | Create an ephemeral node/key; succeed if it didn’t exist. | Write lock record to > 50 % of nodes with a TTL. |
| Hold | Keep the session alive (heartbeat) or rely on TTL renewal. | Refresh TTL before it expires. |
| Release | Delete the node/key or close the session. | Delete the lock record from a majority of nodes. |
The key idea is that the lock’s existence is recorded in a place that all participants trust, and the lock is automatically cleared if the holder crashes (through session expiration or TTL).
Trade‑offs to Discuss
When you talk about distributed locks, interviewers often probe your awareness of the practical compromises:
- Latency vs. safety – Centralised services add a round‑trip for every lock operation, which can be costly in high‑throughput scenarios. Consensus‑based locks spread the load but require multiple round‑trips to achieve a quorum.
- Failure handling – If the coordinator crashes, locks may be orphaned until the service recovers. TTL‑based approaches mitigate this but introduce the risk of premature expiry if a node experiences a temporary pause.
- Fairness – Some implementations provide FIFO ordering (e.g., Zookeeper’s sequential nodes), while others grant the lock to the fastest requester, which can lead to starvation.
- Complexity – Implementing a correct consensus‑based lock is non‑trivial; bugs can cause split‑brain scenarios where two nodes think they hold the lock simultaneously.
- Scalability – Centralised services can become bottlenecks under heavy contention; sharding locks across multiple coordinators can alleviate pressure but adds coordination overhead.
Mentioning these points shows you understand the engineering decisions behind choosing a lock strategy.
Concrete Example: Leader Election in a Microservice Cluster
Suppose you have a set of stateless microservice instances that need to perform a periodic cleanup job (e.g., deleting old logs) without overlapping. You can use a distributed lock to ensure only one instance runs the job at a time.
- Choose a lock provider – Use etcd because it already stores configuration data for the cluster.
- Acquire the lock – Each instance attempts to write a key
cleanup-lockwith a short TTL (e.g., 30 seconds) to a majority of etcd nodes. - Run the job – The instance that succeeds refreshes the TTL every few seconds while the job runs.
- Release – After the job finishes, it deletes the key, allowing the next instance to acquire the lock.
- Failure case – If the instance crashes, the TTL expires, and another instance can acquire the lock automatically.
This pattern illustrates the typical flow and highlights why TTLs are useful for safety.
Common Interview Follow‑up Questions
Interviewers will often dig deeper. Be ready with concise answers to these typical prompts:
"What problems can arise if the lock’s TTL is too short?" A short TTL can cause the lock to expire while the holder is still working, leading to two nodes performing the same critical section simultaneously. The remedy is to choose a TTL comfortably longer than the worst‑case execution time and refresh it periodically.
"How do you avoid deadlocks in a distributed system?" Deadlocks usually stem from acquiring multiple locks in different orders. The safe approach is to enforce a global ordering of lock keys or to use a single lock for the entire critical section when possible.
"Why not just use a database row lock?" Row locks are limited to the database’s transaction scope and require a persistent connection. Distributed locks work across heterogeneous services that may not share a common database, and they survive process restarts because the lock state lives in a dedicated coordination service.
"What is the Redlock algorithm and why is it controversial?" Redlock is a Redis‑based approach that writes the lock to multiple independent Redis instances and requires a majority to succeed. Critics point out that network partitions can still lead to split‑brain scenarios, so it’s recommended only when you can guarantee low latency and strong clock synchronization.
"How would you test a distributed lock implementation?" Simulate node failures, network latency spikes, and clock drift. Verify that only one node can hold the lock at any moment, that the lock recovers after a crash, and that performance meets latency expectations under load.
60‑Second Spoken Answer
"A distributed lock is a coordination primitive that guarantees exclusive access to a shared resource across multiple machines. The most common pattern uses a central service like Zookeeper: a node creates an ephemeral key, and the service grants the lock to the first creator while rejecting others until the key is deleted or the session expires. Alternatively, consensus‑based systems such as etcd write a lock record to a majority of nodes with a TTL; the holder must keep refreshing the TTL to retain ownership. The trade‑offs involve latency—centralised locks add a round‑trip, consensus locks need multiple round‑trips—failure handling, where TTLs protect against crashes but can cause premature expiry, and fairness, since some implementations provide FIFO ordering while others do not. A concrete example is using etcd to elect a leader that runs a periodic cleanup job: each instance tries to write a
cleanup-lockkey with a 30‑second TTL, the winner refreshes the TTL while it works, and if it crashes the TTL expires, allowing another instance to take over. Interviewers often follow up with questions about TTL sizing, deadlock avoidance, and testing strategies."
How to Practice This
- Write the answer on paper – Draft the 60‑second version without looking at notes. Trim any filler until you hit the target length.
- Record yourself – Use a voice recorder or Call Assistant’s practice mode to capture the answer. Listen back for pacing and filler words.
- Simulate follow‑ups – Have a colleague ask the common interview questions listed above. Answer aloud, referencing your resume where you actually used a lock (e.g., “When I built the log‑cleanup service at XYZ, I used etcd…”) to keep the story grounded.
FAQ
- What is the difference between a mutex and a distributed lock? A mutex protects a resource within a single process or machine, while a distributed lock extends that guarantee across multiple machines that communicate over a network.
- Can I implement a distributed lock with only a relational database? Yes, by using a row with a unique constraint and a timestamp, but this approach ties the lock to the database’s transaction model and can be less resilient to network partitions.
- Why is a TTL important for safety? The TTL ensures that if the lock holder crashes or loses connectivity, the lock will eventually expire, preventing permanent deadlock.
- When should I avoid using a distributed lock? If the operation can be made idempotent or if you can redesign the system to use eventual consistency, you may eliminate the need for a lock and reduce complexity.
Frequently asked questions
What is the difference between a mutex and a distributed lock?
A mutex synchronises threads within a single process or host, whereas a distributed lock coordinates exclusive access across multiple machines that communicate over a network.
Can I implement a distributed lock with only a relational database?
You can, by using a uniquely indexed row with a timestamp and checking that the timestamp is recent, but this ties the lock to the database’s transaction semantics and can be vulnerable to network partitions.
Why is a TTL important for safety?
TTL automatically clears the lock if the holder crashes or loses connectivity, preventing permanent deadlock and ensuring another node can eventually acquire the lock.
When should I avoid using a distributed lock?
If the operation can be made idempotent, or if you can redesign the workflow to rely on eventual consistency, you may remove the need for a lock and simplify the system.
#concept#distributed locks#systems design#interview prep#macOS