When an interviewer asks you to design a Content Delivery Network (CDN), they are looking for three things: a clear set of requirements, a sensible decomposition of the system, and an ability to argue about trade‑offs. The conversation usually starts with a high‑level sketch, then narrows down to the parts that are hardest to get right – cache consistency, routing, and fault tolerance. Below is a practical walkthrough you can follow in a real interview.

1. Clarify the scope and requirements

Functional requirements

  • Serve static assets (images, CSS, JS, video) to end users.
  • Provide a simple API for content publishers to upload, purge, and query assets.
  • Support versioning or cache‑busting so that updates propagate quickly.

Non‑functional requirements

  • Low latency (edge‑to‑client round‑trip should be comparable to a local cache).
  • High availability (service should survive node or region failures).
  • Strong consistency for purge operations (once a purge is issued, stale copies must disappear quickly).
  • Scalability to handle traffic spikes without manual re‑configuration.
  • Cost‑effectiveness – use cheaper storage for rarely accessed objects.

Ask the interviewer to prioritize any of these. For example, “Do we need strong consistency for purges, or is eventual consistency acceptable?” The answer will guide later design choices.

2. Identify core entities and API surface

EntityResponsibilityTypical operations
EdgeNodeStores cached copies close to users.GET /asset/:id, PUT /asset/:id, PURGE /asset/:id
OriginServerAuthoritative source of truth.GET /asset/:id, PUT /asset/:id
RouterMaps a request to the nearest healthy EdgeNode.resolve(clientIP, assetID)
CacheControllerDecides what to store, eviction policy, and when to invalidate.store(asset), invalidate(assetID)
ManagementAPIAllows publishers to upload, delete, or purge content.POST /upload, DELETE /asset/:id, POST /purge

The API can be expressed as a set of REST‑style endpoints, but the exact protocol (HTTP, gRPC, etc.) is not the focus – what matters is that the interface cleanly separates the control plane (management) from the data plane (serving).

3. High‑level architecture diagram (text)

Client --> LoadBalancer --> Router --> EdgeNode (cache) <---> OriginServer
                      ^                |
                      |                v
                  ManagementAPI   CacheController
  • LoadBalancer distributes incoming client connections across regional routers.
  • Router selects the optimal EdgeNode based on latency, health, and capacity.
  • EdgeNode serves cached content and falls back to the OriginServer on a miss.
  • CacheController runs on each EdgeNode, handling eviction (LRU, LFU) and purge propagation.
  • ManagementAPI lives on the control plane and talks to all EdgeNodes to push new assets or purge stale ones.

4. Deep dive: Hard parts

4.1 Cache consistency and purge propagation

Purging is the most interview‑friendly challenge. You need a mechanism that guarantees that, after a purge request, no edge will serve the old version. Common approaches:

  • Push‑based invalidation: The ManagementAPI sends a purge message to every EdgeNode via a pub/sub system. EdgeNodes delete the entry immediately.
  • Versioned URLs: Publishers embed a version hash in the URL (e.g., style.v3.css). This sidesteps invalidation but adds complexity for cache‑friendly URLs.
  • Stale‑while‑revalidate: EdgeNode serves the stale copy for a short grace period while fetching the fresh version in the background.

When discussing these, weigh latency (push is fast but requires reliable messaging) against scalability (versioned URLs scale well because they avoid coordination).

4.2 Routing and latency optimization

A router must map a client IP to the nearest EdgeNode. Typical techniques:

  • Geolocation databases (e.g., MaxMind) to map IP ranges to regions.
  • Latency measurements – periodic pings from EdgeNodes to a set of test points, building a latency matrix.
  • Anycast IP – announce the same IP from many locations; the network routes the client to the nearest point automatically.

Explain why you might start with a simple geolocation lookup and then evolve to anycast for production‑grade latency.

4.3 Load balancing and fault tolerance

EdgeNodes should be stateless with respect to request handling; state lives in the cache and the underlying storage. To survive a node failure:

  • Replication – store each cached object on at least two EdgeNodes in the same region.
  • Health checks – routers stop sending traffic to nodes that fail health probes.
  • Graceful degradation – if all replicas are down, the router falls back to the OriginServer.

Mention that replication increases storage cost, so you may choose a replication factor based on the asset’s popularity.

4.4 Storage tiering and cost control

Most CDNs separate hot and cold storage:

  • Hot tier – SSD‑backed caches on EdgeNodes for frequently accessed assets.
  • Cold tier – Object storage (e.g., S3‑compatible) at the origin for infrequently accessed objects.

EdgeNodes can evict items based on LRU, but they may also keep a small “warm” buffer of recently purged items to avoid a thundering‑herd problem when many clients request the same newly‑purged asset.

5. Trade‑off discussion

ConcernOptionProsCons
Purging speedPush invalidationImmediate consistencyRequires reliable messaging, higher control‑plane traffic
Versioned URLsNo coordination neededRequires publisher changes, breaks existing caches
LatencyAnycast routingNetwork does the heavy lifting, low RTTComplex to set up, needs BGP announcements
Geolocation lookupSimple to implement, deterministicCoarse granularity, may route sub‑optimally
Storage costReplicate all objectsHigh availability, fast failoverExpensive, wasteful for rarely accessed content
Selective replicationSaves cost, still protects hot assetsRequires popularity tracking, adds logic

When the interviewer asks “What would you change if the traffic grew tenfold?”, you can point to scaling the pub/sub layer, adding more EdgeNode clusters, and moving from a monolithic router to a distributed hash‑ring.

6. Typical follow‑up questions

  • How would you handle dynamic content (e.g., personalized HTML)?
    • Answer: Use edge‑side scripting (e.g., Cloudflare Workers) to assemble static fragments with user‑specific data, keeping the cacheable portion unchanged.
  • What metrics would you monitor?
    • Answer: Cache hit ratio, purge latency, edge node health, request latency per region, and storage utilization.
  • How do you protect against a cache‑poisoning attack?
    • Answer: Validate origin responses, enforce content‑type checks, and limit the size of objects that can be cached.
  • If the origin is slow, how do you keep latency low?
    • Answer: Pre‑warm caches for popular assets, use background fetches, and fall back to a secondary origin.

7. Sample answer snippet (45‑90 seconds)

“A CDN consists of edge nodes that cache static assets close to users, a routing layer that directs requests to the nearest healthy edge, and a control plane that lets publishers upload and purge content. The core API includes GET /asset/:id for serving, POST /upload for publishing, and POST /purge for invalidation. For consistency, I’d use a push‑based purge over a pub/sub channel so that every edge deletes the stale copy immediately. Routing would start with a geolocation lookup and could evolve to anycast for sub‑millisecond latency. To keep costs in check, hot objects stay on SSD‑backed edge caches while cold objects reside in object storage at the origin, with replication only for the most popular items. Monitoring would focus on cache hit ratio, purge latency, and edge health, and I’d scale the system by adding more edge clusters and sharding the router.”

8. How to practice this

  1. Sketch the diagram on paper – start with the text diagram above, then add more detail each time you rehearse.
  2. Run a mock interview – use Call Assistant to record your answer, then listen back to ensure you stay on topic and reference your own resume where appropriate.
  3. Deep‑dive one hard part per session – pick cache invalidation, routing, or replication and write a short explainer, then rehearse it until you can discuss trade‑offs fluently.

FAQ

  1. Q: Do I need to implement a real CDN to answer this question? A: No. The interview is about architectural reasoning, not code. Focus on components, interfaces, and trade‑offs.
  2. Q: How much detail should I give about the underlying network? A: Mention high‑level concepts like anycast and BGP if asked, but avoid deep routing protocol specifics unless the interviewer probes.
  3. Q: Should I talk about security features like TLS termination? A: Briefly note that edge nodes terminate TLS and forward validated requests, then move on unless security is a primary focus.
  4. Q: What if the interviewer asks for a quantitative estimate? A: Use ranges (“typically a few hundred megabytes per edge node for hot cache”) and explain the factors that would affect the numbers.

Frequently asked questions

Do I need to implement a real CDN to answer this question?

No. The interview is about architectural reasoning, not code. Focus on components, interfaces, and trade‑offs.

How much detail should I give about the underlying network?

Mention high‑level concepts like anycast and BGP if asked, but avoid deep routing protocol specifics unless the interviewer probes.

Should I talk about security features like TLS termination?

Briefly note that edge nodes terminate TLS and forward validated requests, then move on unless security is a primary focus.

What if the interviewer asks for a quantitative estimate?

Use ranges (“typically a few hundred megabytes per edge node for hot cache”) and explain the factors that would affect the numbers.

#system design#cdn#architecture#interview#scalability#a CDN