When an interviewer asks about circuit breakers, they want to see that you understand both the idea and the practical implications. A good answer is short, concrete, and tied to something you actually built.
One‑Sentence Definition
A circuit breaker is a pattern that temporarily stops calls to a downstream service when it is unhealthy, preventing cascading failures and giving the service time to recover.
How It Works
- Monitoring – The client tracks metrics such as error rate, timeout count, or response latency for each downstream endpoint.
- Threshold – When a metric exceeds a configurable limit (e.g., 5 % errors over a 30‑second window), the breaker opens.
- Open State – While open, the client short‑circuits calls and returns an error or fallback immediately.
- Half‑Open Test – After a cool‑down period, a small number of requests are allowed through. If they succeed, the breaker closes; otherwise it stays open.
| State | Behavior | When it changes |
|---|---|---|
| Closed | Calls pass through; metrics are collected. | Normal operation. |
| Open | Calls are blocked; fallback is returned. | Error rate > threshold. |
| Half‑Open | A limited probe traffic is allowed. | Cool‑down timer expires. |
Trade‑offs to Discuss
- Latency vs. Safety – A longer cool‑down reduces load on a recovering service but increases wait time for callers.
- False Positives – Aggressive thresholds can open the circuit for transient spikes, hurting user experience.
- Complexity – Adding a breaker introduces state that must be persisted or shared across instances for consistency.
- Observability – You need good metrics and alerts; otherwise you may not know when the breaker is misbehaving.
Concrete Example from a Real Project
Situation: In a microservice that aggregates pricing data from three external vendors, one vendor started returning 500‑level errors during a traffic surge. Implementation: We added a circuit breaker around the HTTP client for that vendor using the open‑source library Resilience4j. The breaker opened after 3 % errors in a 10‑second window and stayed open for 30 seconds. Result: The overall API latency dropped from 2.4 s to 1.1 s, and the failure rate for end users fell by more than half because we fell back to cached prices instead of propagating the vendor’s errors.
Typical Interviewer Questions
- Why not just use retries? – Retries can amplify load on a failing service; a breaker stops the traffic entirely until the service stabilises.
- How do you choose the thresholds? – Start with production metrics (error rate, latency) and tune based on the service’s SLA. You can also use a percentile‑based approach for latency.
- What’s the difference between a circuit breaker and a bulkhead? – A bulkhead limits concurrency to protect resources, while a breaker stops calls based on health signals.
- How do you handle state in a distributed system? – You can keep the state locally per instance (simpler) or use a shared store (e.g., Redis) for consistent behavior across replicas.
- What fallback strategies are common? – Returning cached data, a default value, or a graceful error message.
60‑Second Spoken Version
"A circuit breaker is a safety valve for service‑to‑service calls. It watches error rates or latency, and if those cross a threshold it opens the circuit—meaning the client stops sending requests and returns a fallback right away. After a cool‑down period it goes half‑open, lets a few calls through, and if they succeed it closes again. The main trade‑off is between protecting the system and adding latency for callers; you need good metrics to set thresholds that avoid false positives. In my last project, we wrapped an unreliable pricing vendor with a breaker that opened after 3 % errors in ten seconds and stayed open for thirty seconds. That cut our API latency in half and reduced user‑visible failures dramatically. When interviewers ask about it, they usually want to know why you’d choose a breaker over retries, how you pick thresholds, and what fallback you provide."
How to Practice This
- Write the answer on paper – Keep it under 150 words and include the definition, mechanism, trade‑off, and a brief example.
- Record yourself – Use a phone or a voice recorder and aim for 45‑90 seconds. Listen back and trim any filler.
- Simulate follow‑up questions – Have a friend ask the typical interview questions listed above, and answer them while keeping the conversation grounded in your résumé. Use Call Assistant to capture the dialogue and get instant feedback on phrasing and relevance.
Frequently asked questions
When should I use a circuit breaker instead of retries?
Use a breaker when the downstream service is consistently failing or overloaded. Retries can increase load and worsen the problem, while a breaker stops traffic until the service recovers.
What is the difference between open and half‑open states?
Open blocks all calls and returns a fallback. Half‑open allows a limited number of test calls after a cool‑down; success closes the circuit, failure keeps it open.
How do I choose the error‑rate threshold?
Start with production metrics—look at typical error rates and latency percentiles. Set the threshold a bit higher than normal spikes, then adjust based on observed false positives.
Can circuit breakers be shared across service instances?
Yes, you can store the breaker state in a distributed cache like Redis, but many teams keep it local for simplicity if the service is stateless and the impact of a brief inconsistency is acceptable.
#concept#circuit breakers#interview#microservices#resilience