When an interview asks you to design a recommendation system, the first thing you should do is pause and ask a few clarifying questions. You want to know the scope of the product, the audience, and the performance expectations. A typical prompt might be: "Design a system that recommends movies to users on a streaming platform." From there you can structure the conversation around functional requirements, non‑functional constraints, the core data model, the API surface, a high‑level architecture, and the deep‑dive topics that interviewers love to probe.

Functional Requirements

  • Personalized list – Return a ranked list of items tailored to each user.
  • Real‑time updates – Incorporate a user’s latest actions (e.g., a recent watch) into the next recommendation.
  • Diverse content – Avoid showing the same genre repeatedly; surface novel items.
  • Explainability (optional) – Provide a short reason why an item was suggested.
  • A/B testing hook – Allow the front‑end to tag a request with an experiment bucket.

Non‑Functional Requirements

RequirementTypical TargetWhy It Matters
Latency< 200 ms for a 20‑item listUsers expect instant suggestions when browsing.
ThroughputHundreds of requests per second per regionTraffic spikes during new releases.
Availability99.9 %+ uptimeDowntime directly hurts engagement metrics.
Data freshnessMinutes for batch updates, seconds for streaming updatesCold‑start users need quick signals; hot users need recency.
ScalabilityHorizontal scaling of storage and computeGrowth in users and catalog is inevitable.

Core Entities and Data Model

  • User – user_id, profile attributes, interaction history.
  • Item – item_id, metadata (title, genre, tags), content embeddings.
  • Interaction – user_id, item_id, event_type (view, like, rating), timestamp.
  • Feature vectors – Pre‑computed embeddings for items (e.g., from a neural model) and optionally for users.

These entities live in two main stores:

  1. Cold store (e.g., a columnar data warehouse) for batch‑computed signals.
  2. Hot store (e.g., a key‑value cache) for real‑time interaction logs.

API Surface

GET /recommendations?user_id={uid}&limit={n}&experiment={exp_id}

Response (JSON):

{
  "user_id": "u123",
  "recommendations": [
    {"item_id": "m456", "score": 0.92, "reason": "Because you liked similar thrillers"},
    {"item_id": "m789", "score": 0.87}
  ]
}

A simple POST endpoint can be used for logging feedback:

POST /feedback
{ "user_id": "u123", "item_id": "m456", "event": "like" }

The API stays thin; the heavy lifting happens inside the service layer.

High‑Level Design

+-------------------+          +-------------------+          +-------------------+
|   Front‑end UI    |  HTTP    |  Recommendation   |  RPC/Batch|  Feature Store   |
| (mobile/web)      |<-------->|   Service Layer   |<-------->| (warehouse)      |
+-------------------+          +-------------------+          +-------------------+
                                   |   ^   |
                                   |   |   |   (real‑time stream)
                                   v   |   v
                           +-------------------+          +-------------------+
                           |   Real‑time Cache |          |   Batch Jobs      |
                           | (e.g., Redis)     |          | (Spark/Beam)      |
                           +-------------------+          +-------------------+

Key components

  • API Gateway – Handles authentication, rate‑limiting, and routes to the recommendation service.
  • Recommendation Service – Orchestrates three pipelines:
    1. Candidate Generation – Fast, coarse filtering (e.g., top‑k popular items, category‑based filters).
    2. Scoring / Ranking – Machine‑learning model that takes user and item vectors, returns a relevance score.
    3. Post‑processing – Re‑rank for diversity, apply business rules, attach explanations.
  • Feature Store – Serves pre‑computed embeddings and metadata; refreshed nightly or hourly.
  • Real‑time Interaction Layer – Consumes a stream of user events (Kafka/ Pulsar) and updates a hot cache.
  • Offline Training Pipeline – Periodically retrains collaborative‑filtering models (matrix factorization, deep factorization machines) using the cold store.

Deep Dive: Candidate Generation

Generating candidates efficiently is often the bottleneck. A naïve approach—scanning the entire catalog for each request—doesn’t scale. Instead, use a two‑tier strategy:

  1. Rule‑based filters – Limit to the user’s preferred genres, language, and age rating.
  2. Approximate Nearest Neighbor (ANN) index – Retrieve the top‑k items whose embeddings are closest to the user vector. Libraries such as Faiss or ScaNN provide sub‑millisecond latency even for millions of vectors.

Trade‑offs

  • Recall vs. latency – A larger candidate pool improves recall but adds latency. Tune the pool size (e.g., 500‑1000) based on the latency budget.
  • Cold‑start users – For brand‑new users, fall back to popularity or demographic‑based heuristics until enough interaction data accumulates.

Deep Dive: Real‑time Personalization

Static models capture long‑term preferences but miss the latest user actions. To keep recommendations fresh:

  • Event stream – Ingest clicks, watches, and likes into a message broker.
  • Feature update – Incrementally adjust a user’s short‑term vector (e.g., using a decay‑weighted sum of recent item embeddings).
  • Cache lookup – Store the updated vector in a fast cache; the ranking stage reads from there.

Challenges

  • State consistency – Ensure that the cache reflects the latest events without causing race conditions.
  • Memory pressure – Hot caches can grow large; employ TTLs and eviction policies based on activity.

Trade‑offs and Alternatives

AspectApproach A: Pure Collaborative FilteringApproach B: Hybrid (CF + Content)
Cold‑startPoor – needs many interactionsGood – content embeddings provide signals
ScalabilitySimple matrix factorization scales with GPU clustersSlightly more complex but still horizontally scalable
ExplainabilityLow – scores are opaqueHigher – content tags can be used for reasons
MaintenancePeriodic retraining onlyRequires both batch and streaming pipelines

Choosing a hybrid design is usually safer because it covers both cold‑start and scalability concerns. However, it adds engineering overhead: you must maintain two pipelines and keep their outputs in sync.

Typical Follow‑Up Questions

  1. "What if we need to support 10 M users and 1 M items?" – Talk about sharding the feature store by user hash, using distributed ANN indexes, and partitioning the interaction stream.
  2. "How would you handle a sudden surge in traffic for a new release?" – Discuss autoscaling the API layer, warm‑up caches, and a fallback to popularity‑based recommendations.
  3. "Can you make the system explain why an item was recommended?" – Show how to attach the top contributing features (e.g., “Because you watched X”) from the ranking model.
  4. "What if the latency budget drops to 50 ms?" – Suggest pruning the candidate pool, moving more logic to the cache, and pre‑computing top‑k lists per user segment.
  5. "How would you evaluate the quality of your recommendations?" – Mention offline metrics (precision@k, NDCG) and online A/B testing with click‑through rate and dwell time.

Sample Answer (45‑90 seconds)

"Sure, let me walk you through a typical recommendation system. First, we clarify the goal: we need to return a personalized, 20‑item list within 200 ms. The core entities are users, items, and interactions, stored in a cold warehouse for batch features and a hot cache for real‑time events. The API is a simple GET endpoint that takes a user ID and limit.

At a high level, the service has three stages. Candidate generation quickly narrows the catalog using rule‑based filters and an ANN index on item embeddings. Then a ranking model scores each candidate using a hybrid of collaborative‑filtering vectors and content features. Finally, we post‑process the list for diversity and add a short reason like ‘Because you liked similar thrillers.’

For freshness, we consume a stream of user actions, update a short‑term user vector in a Redis cache, and read that vector at ranking time. This keeps latency low while still reacting to the latest clicks. Trade‑offs include balancing recall with latency—larger candidate pools improve recall but add latency—and handling cold‑start users by falling back to popularity or demographic heuristics.

If you ask about scaling to millions of users, we shard the feature store by user hash, distribute the ANN index across nodes, and autoscale the API tier. For explainability, we surface the top contributing features from the model. That’s the gist of the design.”

How to practice this

  1. Mock a full interview – Use a timer, ask yourself the clarifying questions, and deliver the answer within 2 minutes.
  2. Sketch the diagram on paper – Practice drawing the components and their connections without digital aids.
  3. Iterate on trade‑offs – Pick a design decision (e.g., candidate pool size) and argue both pros and cons, then switch perspective and defend the opposite choice.

Call Assistant can help you rehearse the spoken version of this answer, keeping you on track and grounding your story in the projects listed on your resume.

Frequently asked questions

What are the essential components of a recommendation system design?

You need candidate generation, a ranking model, a post‑processing layer for diversity and business rules, a feature store for embeddings, a real‑time interaction cache, and an API gateway to expose the service.

How do you handle cold‑start users in a recommendation system?

Use fallback strategies such as popularity‑based lists, demographic heuristics, or content‑based similarity until enough interaction data is collected to power collaborative filtering.

What trade‑offs should I discuss when asked about latency vs. recall?

Explain that a larger candidate pool improves recall but increases latency, so you tune the pool size based on the latency budget and may use caching or pre‑computed top‑k lists to mitigate the impact.

Why is a hybrid approach (collaborative + content) often recommended?

It covers both cold‑start scenarios and provides richer signals for ranking, while still scaling horizontally. The downside is added engineering complexity for maintaining two pipelines.

#system design#recommendation system#architecture#interview prep#scalability#a recommendation system