Load balancing is a core piece of modern infrastructure, but interviewers often expect a crisp, layered answer that shows you understand both the concept and its practical implications.

One‑Sentence Definition

A load balancer is a system that distributes incoming client requests across a pool of backend servers to maximize throughput, minimize latency, and provide fault tolerance.

How It Works: The Main Mechanisms

1. DNS Round‑Robin

  • The simplest form; the DNS server returns multiple IP addresses in a rotating order.
  • No health checking – if a server goes down, clients may still be directed to it until the DNS cache expires.

2. Layer‑4 (Transport) Balancing

  • Operates on TCP/UDP ports.
  • Uses algorithms like Round‑Robin, Least Connections, or Weighted distribution.
  • Can perform health checks by opening a TCP socket; if the handshake fails, the server is removed from the pool.
  • Stateless – the balancer does not inspect the payload, making it fast and easy to scale.

3. Layer‑7 (Application) Balancing

  • Works at the HTTP/HTTPS level.
  • Can route based on URL path, host header, cookies, or request content.
  • Enables sticky sessions (session affinity) by inserting a cookie that ties a client to a particular backend.
  • Often includes SSL termination, caching, and request rewriting.

4. Global Load Balancing (GSLB)

  • Distributes traffic across geographically dispersed data centers.
  • Uses latency measurements, health checks, and DNS tricks to send users to the nearest healthy region.

Trade‑offs to Discuss

AspectLayer‑4Layer‑7DNS Round‑Robin
Latency overheadLow (just TCP)Higher (HTTP parsing)Minimal (DNS only)
FlexibilityLimited (port only)High (content‑aware)None (no health checks)
ScalabilityVery highHigh, but CPU‑intensiveVery high
State managementStatelessCan be stateful (sticky)Stateless
CostLow to moderateHigher (more CPU)Negligible

When to choose each:

  • Use DNS round‑robin for simple, low‑traffic services where occasional failures are acceptable.
  • Choose Layer‑4 for high‑throughput, latency‑sensitive workloads that don’t need request‑level routing.
  • Opt for Layer‑7 when you need content‑based routing, SSL termination, or session affinity.
  • Deploy GSLB for multi‑region redundancy and performance.

Concrete Example

Imagine an e‑commerce site that experiences spikes during a flash sale. The architecture might look like this:

  1. DNS points shop.example.com to a single anycast IP.
  2. Edge router forwards traffic to an L7 load balancer (e.g., NGINX or a cloud‑native balancer).
  3. The balancer terminates TLS, inspects the request path, and routes /api/* to a pool of API servers and /static/* to a CDN.
  4. For the API pool, the balancer uses Least Connections with health checks on /healthz.
  5. If a server fails, the health check removes it, and traffic automatically shifts to the remaining nodes, keeping the sale alive.

Questions Interviewers Frequently Ask

  1. “Can you describe the difference between round‑robin and least‑connections?” – Highlight that round‑robin cycles through servers regardless of load, while least‑connections directs traffic to the server with the fewest active connections, which can better handle uneven request sizes.
  2. “How do you handle sticky sessions?” – Explain using a balancer‑generated cookie or a consistent‑hashing algorithm that maps a client’s identifier to a specific backend.
  3. “What happens if the load balancer itself fails?” – Discuss high‑availability setups: active‑passive pairs, virtual IP failover, or using a cloud provider’s managed service that automatically routes around failures.
  4. “When would you prefer DNS‑level load balancing over an L4/L7 balancer?” – Mention cost considerations, simplicity, or when you need to distribute traffic across regions without a dedicated balancer.
  5. “How do you measure the effectiveness of a load‑balancing strategy?” – Talk about metrics such as request latency, error rate, server CPU/memory utilization, and connection counts, often visualized in dashboards.

60‑Second Spoken Answer (Ready for the Interview)

“A load balancer spreads incoming requests across multiple backend servers to keep the system responsive and resilient. The simplest method is DNS round‑robin, which just rotates IPs, but it lacks health checks. More robust solutions work at layer‑4, distributing traffic based on TCP/UDP ports using algorithms like round‑robin or least connections, and they can quickly drop unhealthy nodes. Layer‑7 balancers operate at the HTTP level, allowing routing by URL, host header, or cookies, and they can terminate SSL and provide sticky sessions. The trade‑off is that layer‑7 adds processing overhead but gives you fine‑grained control. In practice, I’d use DNS for global distribution, a layer‑4 balancer for high‑throughput services, and a layer‑7 balancer when I need content‑aware routing or session affinity. Monitoring latency, error rates, and connection counts tells you whether the strategy is working.”

How to Practice This

  1. Record yourself answering the 60‑second version and listen for filler words; tighten the phrasing until it fits comfortably under a minute.
  2. Use Call Assistant to simulate follow‑up questions; let it capture your resume points and keep the conversation on topic while you refine the technical story.
  3. Build a mini lab with a cheap cloud instance or local VMs: set up NGINX as a layer‑7 balancer and generate traffic with hey or wrk to observe how different algorithms affect latency and error rates.

FAQ

  • What is the main benefit of a load balancer? It improves availability by routing around failed servers and boosts performance by spreading work across multiple machines.
  • Is DNS round‑robin a true load balancer? It provides basic distribution but lacks health checks, so it’s not a complete solution for high‑availability services.
  • When should I use sticky sessions? When your application stores state locally on a server (e.g., shopping cart) and you need the same client to hit that server for the duration of the session.
  • Can a single load balancer become a bottleneck? Yes; that’s why production setups often run multiple balancers in active‑active mode or rely on managed services that automatically scale.

Frequently asked questions

What is the main benefit of a load balancer?

It improves availability by routing around failed servers and boosts performance by spreading work across multiple machines.

Is DNS round-robin a true load balancer?

It provides basic distribution but lacks health checks, so it’s not a complete solution for high-availability services.

When should I use sticky sessions?

When your application stores state locally on a server (e.g., a shopping cart) and you need the same client to hit that server for the duration of the session.

Can a single load balancer become a bottleneck?

Yes; production setups often run multiple balancers in active‑active mode or use managed services that automatically scale to avoid a single point of congestion.

#concept#load balancing#interview#technical#systems