When an interview asks you to design an API gateway, the goal is to see how you balance business needs with engineering constraints. You need to demonstrate that you can turn vague product goals into a concrete, maintainable system. Below is a walkthrough you can use as a checklist while you’re on the whiteboard.
1. Clarify the Scope
Start by asking the interviewer a few clarifying questions. This shows that you care about the right problem, not just a generic solution.
- Who are the callers? Internal services, external partners, or both?
- What kind of APIs are we exposing? REST, gRPC, GraphQL, or a mix?
- Do we need to support versioning? How often do new versions roll out?
- What are the SLAs? Typical latency targets and availability expectations.
- Any compliance constraints? For example, PCI‑DSS or GDPR may dictate logging and encryption.
By pinning down these variables you avoid over‑engineering and can tailor the design to the most relevant concerns.
2. Functional Requirements
| Requirement | Why It Matters |
|---|---|
| Routing | Directs a request to the correct backend service based on path, method, or headers. |
| Authentication & Authorization | Ensures only allowed callers can hit an endpoint; often integrates with an identity provider. |
| Rate Limiting & Quotas | Protects downstream services from overload and enforces contractual limits. |
| Request Transformation | Handles protocol translation, header injection, or payload reshaping. |
| Observability | Emits metrics, logs, and traces to help operators monitor health and debug issues. |
| Caching | Reduces latency for read‑heavy endpoints and eases load on backends. |
| Health Checks | Provides a quick way to verify that the gateway itself is up and routing correctly. |
These are the core capabilities you’ll need to cover in the design.
3. Non‑Functional Requirements
- Scalability – The gateway must handle traffic spikes without becoming a bottleneck. Horizontal scaling is the usual approach.
- Low Latency – Adding a gateway adds an extra hop; keep processing time under a few milliseconds.
- High Availability – Deploy across multiple zones or regions; use health‑checking and fail‑over.
- Security – TLS termination, secret management, and audit logging are non‑negotiable for production.
- Operability – Deploy‑time configuration changes should be possible without a full restart.
- Observability – Export standardized metrics (e.g., Prometheus) and support distributed tracing.
Mentioning these early signals that you think beyond the code.
4. Core Entities and API
4.1 Data Model
- Route – Defines a pattern (e.g.,
/api/v1/users/*), the target service, and optional rewrite rules. - Policy – Bundles authentication method, rate‑limit settings, and IP allow‑list.
- CacheEntry – Stores a response key, TTL, and optional validation headers.
- HealthProbe – Records the result of a periodic check against a backend.
4.2 Public API (for Ops)
POST /routes {"path":"/orders/*","service":"order-service","methods":["GET","POST"]}
GET /routes → list of routes
PUT /routes/:id {"rateLimit":10}
DELETE /routes/:id
These endpoints let an operator add, modify, or remove routes without redeploying the gateway.
5. High‑Level Architecture
+-------------------+ +-------------------+ +-------------------+
| Client (Web) | ---> | API Gateway | ---> | Service A |
+-------------------+ | (Edge Layer) | +-------------------+
| +-----------+ |
| | Auth/Zauth| |
| +-----------+ |
| +-----------+ |
| | RateLimiter| |
| +-----------+ |
| +-----------+ |
| | Cache Layer| |
| +-----------+ |
+-------------------+
| |
v v
+-------------------+ +-------------------+
| Service B | | Service C |
+-------------------+ +-------------------+
- Edge Layer – Handles TLS termination, quick health checks, and basic routing.
- Auth/Zauth – Calls out to an identity provider (e.g., OIDC) and caches tokens.
- RateLimiter – Implements token‑bucket or leaky‑bucket algorithms per caller.
- Cache Layer – Stores responses in an in‑memory store (e.g., Redis) for read‑heavy endpoints.
- Backend Services – Remain unaware of the gateway; they just receive clean HTTP/gRPC calls.
The diagram emphasizes separation of concerns: each responsibility can be scaled or swapped independently.
6. Deep Dive: Hard Parts
6.1 Routing Performance
Routing is the first thing every request hits, so latency matters. A common pattern is to compile route definitions into a trie or radix tree that can be traversed in O(L) time, where L is the length of the path. Keep the data structure immutable and replace it atomically when a new route is added—this avoids locking on the hot path.
6.2 Distributed Rate Limiting
If you run multiple gateway instances, a naïve per‑instance limit can be bypassed by a client spreading requests across instances. Two practical approaches:
- Centralized token bucket – Store counters in a fast key‑value store (e.g., Redis) and use Lua scripts for atomic updates.
- Consistent hashing – Map a caller’s identifier to a specific instance, giving each instance a deterministic slice of the quota.
Both trade latency for fairness; pick the one that matches the SLA.
6.3 Observability without Overhead
Injecting tracing headers early (e.g., traceparent) and immediately logging request metadata lets you capture end‑to‑end latency. However, logging every request can overwhelm storage. Use sampling (e.g., 1‑in‑100) for high‑volume endpoints and always log error paths.
6.4 Hot‑Reloading Configuration
Operators need to change routes on the fly. Implement a watch on a configuration store (e.g., etcd) that pushes updates to each gateway instance. The instance validates the new config, swaps the routing trie, and continues serving without a pause.
7. Trade‑offs and Alternatives
| Aspect | API Gateway (Edge) | Service Mesh Sidecar |
|---|---|---|
| Latency | Extra hop, but can cache at edge. | Minimal extra hop; traffic stays within mesh. |
| Complexity | Simpler to deploy; central point of failure. | Distributed; each service runs a sidecar, increasing resource usage. |
| Observability | Centralized logs and metrics. | Uniform tracing across services, but harder to aggregate globally. |
| Security | Central TLS termination; easier to enforce policies. | Mutual TLS between services; more granular but requires certificate management. |
Choosing between a gateway and a mesh depends on where you want control: at the edge (gateway) or service‑to‑service (mesh). In many modern stacks, both coexist—gateway for inbound traffic, mesh for internal communication.
8. Common Follow‑Up Questions
- How would you handle versioning? – Keep separate routes per version (e.g.,
/v1/*vs/v2/*) and allow deprecation flags. - What happens if the authentication provider is down? – Cache validated tokens for a short TTL and fall back to a “deny‑all” mode with alerts.
- How do you ensure zero‑downtime deployments? – Use blue‑green or canary releases for the gateway itself, routing a small percentage of traffic to the new version first.
- Can you support GraphQL alongside REST? – Yes; treat the GraphQL endpoint as a single route that forwards the request to a dedicated GraphQL service, optionally enabling query‑level caching.
- What monitoring dashboards would you expose? – A latency histogram per route, request‑per‑second counters, error rates, and cache hit‑miss ratios.
9. Sample Answer (45‑90 seconds)
"In this design I start by listing the functional pieces: routing, auth, rate limiting, caching, and observability. I’d build a lightweight edge layer that terminates TLS, looks up the request in a compiled trie of routes, and forwards it to the appropriate backend. Auth is delegated to an OIDC provider; tokens are cached for a few minutes to avoid repeated calls. Rate limiting is done centrally using a Redis‑backed token bucket so a client can’t bypass limits by hitting different instances. For read‑heavy endpoints I add a cache layer that stores responses for a configurable TTL. All metrics—latency, request counts, error rates—are exported to Prometheus, and traces are propagated via the W3C
traceparentheader. The configuration lives in etcd; each gateway watches for changes and swaps the routing trie atomically, enabling hot‑reload without downtime. Trade‑offs include a small added hop for inbound traffic, but the central point lets us enforce security and policies consistently. If we needed fine‑grained internal security we could complement this with a service mesh for service‑to‑service calls."
Practicing this answer aloud with Call Assistant can help you keep the flow tight and ensure every claim ties back to a concrete experience on your résumé.
How to practice this
- Sketch the diagram on paper – Do it three times, each with a different focus (routing, rate limiting, observability).
- Record yourself answering – Use Call Assistant to capture the answer and get a concise version that stays within the 90‑second window.
- Mock follow‑ups – Have a colleague ask the common questions above and practice extending your answer without losing structure.
FAQ
Q: Do I need to implement TLS termination inside the gateway? A: Yes, terminating TLS at the edge lets you inspect headers for routing and auth. Internally you can re‑encrypt if downstream services require it.
Q: How much caching is reasonable for an API gateway? A: Cache only idempotent GET responses that are safe to serve stale for a short window. Over‑caching can hide bugs in downstream services.
Q: When would a service mesh be a better choice than a gateway? A: If you need fine‑grained, per‑service security (mutual TLS) and want to offload retries, circuit breaking, and observability to the mesh, it complements a gateway rather than replaces it.
Q: What’s a simple way to test the gateway’s rate‑limiting logic? A: Write a script that fires a burst of requests from a single client ID and verify that the response status switches from 200 to 429 after the configured quota is exceeded.
Frequently asked questions
Do I need to implement TLS termination inside the gateway?
Yes, terminating TLS at the edge lets you inspect headers for routing and auth. Internally you can re‑encrypt if downstream services require it.
How much caching is reasonable for an API gateway?
Cache only idempotent GET responses that are safe to serve stale for a short window. Over‑caching can hide bugs in downstream services.
When would a service mesh be a better choice than a gateway?
If you need fine‑grained, per‑service security (mutual TLS) and want to offload retries, circuit breaking, and observability to the mesh, it complements a gateway rather than replaces it.
What’s a simple way to test the gateway’s rate‑limiting logic?
Write a script that fires a burst of requests from a single client ID and verify that the response status switches from 200 to 429 after the configured quota is exceeded.
#system design#api gateway#architecture#scalability#observability#an API gateway