Load balancing shows up in almost every systems interview. Whether you’re talking to a junior hiring manager or a senior architect, the core ideas stay the same: distribute traffic, keep services healthy, and protect the user experience. Below are the most common questions you’ll hear, a concise spoken answer you can deliver in 45‑90 seconds, and the next‑level follow‑up the interviewer often asks. Use the templates as a starting point, then ground them in your own résumé – that’s where Call Assistant can help you rehearse the story and keep the conversation on track.

1. What is a load balancer and why do we need one?

Answer template "A load balancer sits in front of a pool of servers and decides which instance should handle each incoming request. It spreads traffic to avoid overloading any single node, improves latency by routing to the closest healthy instance, and provides a single entry point for TLS termination and health monitoring. Without it, a spike in traffic could crash a service, and you’d lose the ability to roll out updates without downtime." Typical follow‑up: Can you give an example of a failure you prevented with a load balancer?

2. Differentiate between Layer 4 and Layer 7 load balancing.

Answer template "Layer 4 balances based on transport‑layer information – IP address and port. It’s fast because it forwards packets without inspecting the payload. Layer 7 looks at the application layer, so it can route based on URL path, HTTP headers, or cookies. That lets you do content‑based routing, A/B testing, or sticky sessions, but it adds a bit of processing overhead." Typical follow‑up: When would you choose Layer 7 over Layer 4 despite the extra cost?

3. Explain health checks and why they matter.

Answer template "Health checks are periodic probes – usually HTTP GET or TCP connect – that verify each backend is still able to serve traffic. The balancer removes unhealthy nodes from the pool, preventing client errors. You can configure passive checks (like 5xx responses) and active checks with custom endpoints that report internal metrics such as queue depth." Typical follow‑up: How do you design a health‑check endpoint for a microservice?

4. What are sticky sessions and when should you use them?

Answer template "Sticky sessions, or session affinity, bind a client to a specific backend for the duration of the session. They’re useful when the server stores state locally – for example, in‑memory caches or non‑shared session data. You implement them with cookies, source IP hashing, or a dedicated session token. If your architecture is truly stateless, you should avoid stickiness because it reduces load‑balancing efficiency." Typical follow‑up: How would you migrate from sticky sessions to a stateless design?

5. How does TLS termination work in a load balancer?

Answer template "TLS termination means the balancer decrypts incoming HTTPS traffic and forwards plain HTTP to the backends. This offloads the CPU‑intensive handshake from the servers and lets you centralise certificate management. You can also do TLS passthrough, where the balancer forwards encrypted traffic unchanged, which is required when end‑to‑end encryption is a compliance requirement." Typical follow‑up: What are the security implications of terminating TLS at the balancer?

6. Describe a typical load‑balancing algorithm and its trade‑offs.

Answer template "Round‑robin cycles through servers in order – simple and works well when all nodes have similar capacity. Least‑connections sends traffic to the server with the fewest active connections, which helps when request lengths vary. Weighted variants let you bias traffic toward more powerful instances. Consistent hashing is useful for cache‑coherent workloads because it minimizes key movement when nodes change." Typical follow‑up: Which algorithm would you pick for a latency‑sensitive API and why?

7. How do you handle scaling the load balancer itself?

Answer template "Modern cloud providers let you treat the balancer as a managed service that automatically scales horizontally – for example, using an Elastic Load Balancer that adds capacity as request volume rises. In on‑prem setups you’d deploy a pair of HAProxy or NGINX instances behind a virtual IP using VRRP, then add more nodes behind a DNS‑based round‑robin as needed. Monitoring metrics like CPU, connections, and latency tells you when to add capacity." Typical follow‑up: What metrics would you set alerts on to detect a balancer bottleneck?

8. Explain the concept of “failover” vs “load‑balancing”.

Answer template "Load‑balancing spreads traffic under normal conditions. Failover kicks in when a primary node or entire region becomes unavailable – traffic is redirected to a standby pool. In practice, you configure health checks that trigger failover, and you might use DNS‑based routing or a global traffic manager to shift users to a different data center." Typical follow‑up: How would you test a failover scenario without impacting production users?

9. What are “global load balancers” and when are they appropriate?

Answer template "Global load balancers operate at the DNS or Anycast layer to route users to the nearest geographic region. They’re useful for multi‑region deployments where latency matters, or for regulatory reasons that require data to stay in a specific country. They often combine health checks with geo‑routing policies to keep traffic on the healthiest edge location." Typical follow‑up: Can you describe a situation where a global LB introduced new challenges?

10. How do you secure a load balancer against DDoS attacks?

Answer template "First, you place the balancer behind a network‑level DDoS mitigation service that scrubs traffic. Second, you enable rate limiting and connection throttling on the balancer itself. Third, you enforce strict TLS ciphers and use HTTP/2 to reduce header overhead. Finally, you monitor anomalous traffic patterns and automate black‑listing of offending IP ranges." Typical follow‑up: What would you do if an attacker bypassed the edge protection and hit the balancer directly?

11. Sample Answer Walk‑through

Below is a full‑sentence spoken answer that strings together several of the concepts above. Tailor the specifics (e.g., the cloud provider or the project name) to your own experience.

"In my last role, we migrated a monolithic API gateway to a multi‑region architecture. We placed an AWS Global Accelerator in front of two Application Load Balancers – one in us‑east‑1 and another in eu‑central‑1. The ALBs performed TLS termination and used weighted least‑connections to balance traffic across a pool of EC2 instances. Health checks hit a /healthz endpoint that reported queue depth, so the balancers could automatically drain a node that started lagging. For sessions that required affinity – like a shopping cart stored in Redis – we enabled cookie‑based stickiness, but we later refactored the cart service to be stateless, which let us switch to pure round‑robin and improve throughput by about 20 %. When a regional outage occurred, the Global Accelerator automatically rerouted traffic to the healthy region, and we verified the failover by pulling down the primary ALB in a controlled test."

12. Quick Comparison Table

FeatureLayer 4 LBLayer 7 LB
Decision BasisIP/PortURL, Headers, Cookies
PerformanceVery high (packet‑level)Slightly lower (application parsing)
Content‑Based RoutingNoYes
SSL/TLS TerminationUsually passthroughNative termination
Use CasesSimple TCP services, high‑throughput microservicesHTTP APIs, A/B testing, sticky sessions

13. How to practice this

How to practice this

  1. Record yourself – Use Call Assistant to capture a mock interview. Answer a question, then listen back and trim any filler or jargon.
  2. Add a follow‑up – After each answer, have a friend ask a deeper question (e.g., “What metric would you monitor for failover?”). Practice pivoting while keeping the story anchored to your résumé.
  3. Iterate with variations – Swap out the technology stack (NGINX → HAProxy, AWS → Azure) to ensure you understand the concepts, not just the product names.

FAQ

  • Q: What’s the difference between health checks and readiness probes? A: Health checks are performed by the load balancer to decide if a node should receive traffic. Readiness probes are used by orchestration platforms (like Kubernetes) to decide if a pod should be added to the service pool. Both aim to prevent unhealthy instances from serving requests, but they operate at different layers.
  • Q: When should I use DNS‑based load balancing versus a dedicated load balancer? A: DNS‑based balancing is cheap and works for coarse‑grained traffic distribution across regions. A dedicated LB is better for fine‑grained routing, health‑check driven failover, and TLS termination.
  • Q: Is sticky session always a bad practice? A: Not necessarily. It’s acceptable when you have a legitimate need for session affinity, such as legacy stateful services. However, it reduces the effectiveness of load distribution and makes scaling harder, so you should aim to eliminate it when possible.
  • Q: How do I measure the effectiveness of my load‑balancing strategy? A: Track latency percentiles, error rates, and connection counts per backend. Compare before/after metrics when you change algorithms or add capacity to see the impact.

Frequently asked questions

What’s the difference between health checks and readiness probes?

Health checks are performed by the load balancer to decide if a node should receive traffic. Readiness probes are used by orchestration platforms (like Kubernetes) to decide if a pod should be added to the service pool. Both aim to prevent unhealthy instances from serving requests, but they operate at different layers.

When should I use DNS‑based load balancing versus a dedicated load balancer?

DNS‑based balancing is cheap and works for coarse‑grained traffic distribution across regions. A dedicated LB is better for fine‑grained routing, health‑check driven failover, and TLS termination.

Is sticky session always a bad practice?

Not necessarily. It’s acceptable when you have a legitimate need for session affinity, such as legacy stateful services. However, it reduces the effectiveness of load distribution and makes scaling harder, so you should aim to eliminate it when possible.

How do I measure the effectiveness of my load‑balancing strategy?

Track latency percentiles, error rates, and connection counts per backend. Compare before/after metrics when you change algorithms or add capacity to see the impact.

#concept questions#load balancing#interview prep#systems design#technical concepts