Reddit’s system design interview is a deep dive into how you would build or evolve a core piece of the platform. The goal is not to see if you can recite a textbook design, but whether you can think through constraints, make sensible trade‑offs, and communicate a coherent architecture that fits Reddit’s community‑driven product.

What the Round Usually Looks Like

  • Duration: 45‑60 minutes, shared screen, no code editor.
  • Format: You receive a high‑level prompt (e.g., "Design a service that ranks comments in real time"). You ask clarifying questions, outline the system, and discuss bottlenecks.
  • Evaluation Rubric: Interviewers typically score on four axes:
    1. Problem Understanding – how well you scope the problem and identify key metrics.
    2. High‑Level Architecture – clarity of components, data flow, and separation of concerns.
    3. Trade‑Off Analysis – ability to discuss latency vs. consistency, cost vs. performance, etc.
    4. Communication – logical flow, use of diagrams, and handling follow‑up questions.
  • Typical Focus Areas: Content ranking, real‑time notifications, moderation pipelines, and data‑driven personalization. Reddit’s architecture is heavily micro‑service oriented, uses a mix of relational and NoSQL stores, and relies on asynchronous processing for scalability.

Common Design Prompts

Below are two prompts that have appeared in public discussions. Numbers are abstracted; focus on the variables and reasoning.

Prompt 1: Real‑Time Comment Ranking Service

Goal: Build a service that ranks comments for a post as new votes arrive, keeping the ranking list fresh for users scrolling the comment thread.

Key Variables

  • P: number of active posts per second (peak load).
  • C: average comments per post.
  • V: vote events per second.
  • R: required latency for an updated ranking (e.g., < 200 ms).

High‑Level Sketch

  1. API Layer – receives vote events (POST /vote).
  2. In‑Memory Score Store – a sharded key‑value store (e.g., Redis) holding current scores per comment ID.
  3. Ranking Engine – a lightweight process that recomputes the top‑K list on each vote, using a priority queue per post.
  4. Persistence – async write‑behind to a durable store (e.g., Cassandra) for audit and replay.
  5. Read Path – GET /post/{id}/comments pulls the top‑K from the in‑memory store, falls back to a batch fetch from the persistent store for deeper pages.

Trade‑Offs

  • Latency vs. Consistency – In‑memory store gives low latency but can lose votes on crash; write‑behind mitigates data loss.
  • Sharding – Partition by post ID to keep hot posts isolated; however, uneven popularity can cause hotspotting.
  • Ranking Algorithm – Simple score = upvotes‑downvotes works for most cases; for more nuanced ranking (e.g., Reddit’s “hot” algorithm), you need a periodic background job to recompute scores.

Prompt 2: Real‑Time Notification System

Goal: Deliver notifications to users when they receive a reply, a mention, or a moderator action, with near‑real‑time delivery on web and mobile.

Key Variables

  • U: active user base receiving notifications per second.
  • N: average notifications per user per day.
  • D: delivery latency budget (e.g., < 500 ms).

High‑Level Sketch

  1. Event Producer – services emit notification events to a message broker (e.g., Kafka).
  2. Notification Service – consumes events, enriches with user preferences, and writes to a per‑user queue (e.g., DynamoDB Streams or Redis Streams).
  3. Push Dispatcher – reads from the queue and pushes to platform‑specific channels (WebSocket for web, APNs/FCM for mobile).
  4. Storage – a durable log (e.g., Kafka) retains events for replay and auditing.
  5. Read API – GET /notifications pulls recent items from the per‑user store, supporting pagination.

Trade‑Offs

  • At‑Least‑Once vs. Exactly‑Once – Using Kafka’s delivery semantics gives at‑least‑once; deduplication logic in the consumer layer handles duplicates.
  • Fan‑out – Directly pushing to each device can overload the dispatcher; batching per user reduces calls.
  • Scalability – Partition the topic by user ID hash to spread load, but ensure ordering per user for correct sequencing.

Preparing Without the Numbers

Because Reddit’s public data varies by region and time, focus on variables rather than exact figures. When you practice, ask yourself:

  • What is the peak traffic I need to support?
  • How does the system behave when a single post becomes viral?
  • Which components can be cached, and where does latency dominate?

A Practical Prep Plan

  1. Build a Reusable Checklist – Keep a short list of the rubric axes (understanding, architecture, trade‑offs, communication). Before each mock interview, run through the checklist.
  2. Sketch, Iterate, Sketch Again – Use paper or a digital whiteboard to draw component diagrams. After each iteration, ask yourself: "What failure mode am I ignoring?" and add a mitigation.
  3. Talk It Out Loud – The biggest gap for many candidates is the transition from mental model to spoken explanation. Tools like Call Assistant let you rehearse your answer, keep the conversation on track, and surface follow‑up questions you might miss.
  4. Focus on Reddit‑Specific Context – Review Reddit’s public architecture blog posts (e.g., their move to micro‑services, use of PostgreSQL for core data, and adoption of Kafka for event streams). Align your design decisions with those patterns.
  5. Mock Interviews with Peer Feedback – Pair with another engineer and rotate roles: one designs, the other probes. Swap after each prompt and compare notes against the rubric.

How to Practice This

  1. Pick a Prompt – Choose one of the example prompts above or a recent Reddit blog post topic.
  2. Run a 45‑minute Mock – Set a timer, ask clarifying questions, and walk through the four rubric axes.
  3. Record and Review – Use Call Assistant to record your spoken answer, then replay to spot unclear sections or missed trade‑offs.

FAQ

Q: Do I need to know the exact tech stack Reddit uses? A: No. Focus on the principles (micro‑services, async pipelines, sharding). Mentioning that Reddit uses PostgreSQL for relational data and Kafka for event streaming shows awareness, but you can propose alternatives if they make sense for the problem.

Q: How deep should I go into database schema design? A: Briefly outline primary entities and access patterns. Detailed schema is rarely required unless the prompt explicitly asks for data modeling.

Q: What if the interviewer pushes me toward a specific technology? A: Acknowledge the suggestion, explain the trade‑offs of adopting it, and keep the conversation anchored to the problem’s constraints.

Q: How can I keep the discussion from drifting into unrelated features? A: Use a simple “scope reminder” – periodically restate the core goal (e.g., "Our focus is delivering a ranked comment list within 200 ms") to keep the design tight.

Frequently asked questions

Do I need to know the exact tech stack Reddit uses?

No. Focus on the principles (micro‑services, async pipelines, sharding). Mentioning that Reddit uses PostgreSQL for relational data and Kafka for event streaming shows awareness, but you can propose alternatives if they make sense for the problem.

How deep should I go into database schema design?

Briefly outline primary entities and access patterns. Detailed schema is rarely required unless the prompt explicitly asks for data modeling.

What if the interviewer pushes me toward a specific technology?

Acknowledge the suggestion, explain the trade‑offs of adopting it, and keep the conversation anchored to the problem’s constraints.

How can I keep the discussion from drifting into unrelated features?

Use a simple “scope reminder” – periodically restate the core goal (e.g., "Our focus is delivering a ranked comment list within 200 ms") to keep the design tight.

#Reddit#system design#interview prep#architecture#tech interview