You walk into a system design interview and the recruiter asks you to design a live comments system – the kind you see under a streaming video or a live article. The goal is to show that you can translate a vague product idea into a concrete architecture, reason about trade‑offs, and keep the conversation focused.

1. Clarify the Problem

Start by asking clarifying questions. Typical points include:

  • Scope – Is it a single‑tenant product (e.g., an internal blog) or a multi‑tenant platform serving many publishers?
  • Latency – How fast must a new comment appear to other viewers? Sub‑second latency is common for live experiences.
  • Throughput – What are the expected peaks? A popular live event can generate thousands of comments per second, while a low‑traffic article may see only a few per minute.
  • Durability – Must comments survive crashes? Are they ever edited or deleted?
  • Moderation – Real‑time profanity filtering? Human review?
  • Feature set – Replies, threading, likes, pinning, or reactions?

These answers shape the rest of the design and give you a chance to demonstrate product sense.

2. Functional Requirements

RequirementWhy it matters
Post a commentUsers submit text, optionally with metadata (user id, timestamp, parent id).
Read a comment streamViewers see a continuously updating list, ordered by time or relevance.
Pagination / Infinite scrollClients load older comments without overwhelming the server.
Edit / DeleteUsers may correct mistakes; moderation may need to remove abusive content.
Moderation hooksAutomatic filters and human review keep the conversation safe.
MetricsTrack comment volume, latency, and error rates for monitoring.

3. Non‑Functional Requirements

  • Low latency – Aim for < 500 ms end‑to‑end for a newly posted comment to appear.
  • Scalability – System should handle traffic spikes (e.g., a celebrity livestream) without degradation.
  • Availability – At least 99.9 % uptime; graceful degradation (e.g., show cached comments) if a component fails.
  • Durability – Persist comments to durable storage; loss of a few seconds may be acceptable but not permanent loss.
  • Consistency – Eventual consistency is fine for read‑after‑write; strong ordering is needed within a single stream.
  • Security – Authenticate users, authorize posting to a given stream, and protect against injection attacks.

4. Core Entities & API

4.1 Data Model (simplified)

Comment {
  id: UUID,
  streamId: UUID,
  userId: UUID,
  parentId: UUID?,   // null for top‑level comment
  body: string,
  createdAt: timestamp,
  editedAt: timestamp?,
  isDeleted: bool,
  reactions: map<string, int>
}

4.2 Public API (REST‑like)

MethodEndpointPurpose
POST/streams/{streamId}/commentsCreate a new comment.
GET/streams/{streamId}/comments?after={cursor}&limit={n}Pull newer comments after a cursor (for live updates).
GET/streams/{streamId}/comments?before={cursor}&limit={n}Load older comments (pagination).
PATCH/comments/{commentId}Edit a comment (if allowed).
DELETE/comments/{commentId}Soft‑delete a comment.

You can also expose a WebSocket or Server‑Sent Events (SSE) endpoint for push‑based delivery, which many interviewers expect for a live system.

5. High‑Level Architecture

+----------------+      +-------------------+      +-------------------+
|   Client UI    |<---->|   API Gateway     |<---->|   Auth Service    |
+----------------+      +-------------------+      +-------------------+
          |                       |                       |
          v                       v                       v
+----------------+      +-------------------+      +-------------------+
|   Write Service|      |   Read Service    |      |   Moderation Svc |
+----------------+      +-------------------+      +-------------------+
          |                       |                       |
          v                       v                       v
+----------------+      +-------------------+      +-------------------+
|   Append‑Only  |      |   In‑Memory Cache |      |   Profanity DB   |
|   Log (Kafka) |      |   (Redis)         |      +-------------------+
+----------------+      +-------------------+                |
          |                       |                        v
          v                       v                +-------------------+
+----------------+      +-------------------+     |   Durable Store   |
|   Worker Pool  |----->|   Fan‑out Service  |---->|   (PostgreSQL)    |
+----------------+      +-------------------+     +-------------------+

5.1 Write Path

  1. API Gateway validates the request and forwards it to the Write Service.
  2. The Write Service writes the comment to an append‑only log (Kafka‑style). This provides durability and a replayable source for downstream consumers.
  3. A Worker Pool consumes the log, runs moderation checks, and writes the sanitized comment to the Durable Store.
  4. The same worker pushes the comment to a Fan‑out Service that publishes to a Pub/Sub channel per streamId.
  5. The Read Service (or push endpoint) receives the message and updates the In‑Memory Cache (e.g., a Redis sorted set keyed by streamId).

5.2 Read Path

  • Polling: Clients call the GET endpoint with a cursor token. The Read Service reads from the cache first, falling back to the DB for older pages.
  • Push: For live updates, the client opens an SSE/WebSocket. The Read Service streams new entries from the Pub/Sub channel directly to the client.

6. Deep Dive: Scaling the Hot Write Path

6.1 Sharding

When many streams are active simultaneously, a single Kafka topic can become a bottleneck. Partition the log by streamId hash, ensuring that all comments for the same stream land in the same partition (preserving order) while spreading load across many brokers.

6.2 Back‑pressure & Rate Limiting

A sudden surge (e.g., a viral moment) can flood the write path. Implement per‑user and per‑stream rate limits at the API Gateway. If the queue length exceeds a threshold, the gateway can return a 429 response, prompting the client to retry with exponential back‑off.

6.3 Caching Strategy

  • Hot streams (active live events) stay in Redis with a TTL of a few minutes after the last comment.
  • Cold streams fall back to the relational store; pagination queries hit the DB directly.
  • Use a sorted set (ZADD) keyed by streamId with the timestamp as the score, enabling O(log N) range queries for recent comments.

7. Trade‑offs & Alternatives

AspectChoiceProsCons
Log vs Direct DB writeAppend‑only logGuarantees ordering, easy replay, decouples producer/consumerAdds latency before DB persistence
Push vs PullSSE/WebSocketNear‑real‑time, reduces client pollingRequires connection management, scaling sockets
In‑Memory cacheRedisFast reads, supports sorted queriesNeeds eviction policy, extra ops cost
StorageRelational DBStrong consistency, flexible queriesMay struggle with very high write rates without sharding
ModerationSynchronous filterImmediate feedback to userIncreases write latency
Asynchronous reviewKeeps path fast, handles complex rulesUsers may see abusive content briefly

Choosing the right combination depends on the product’s tolerance for latency vs. consistency and the expected traffic pattern.

8. Typical Follow‑Up Questions

  1. How would you handle comment ordering across partitions? – Explain that ordering is only required within a single stream, so partitioning by streamId preserves order. For global ordering you could use a globally increasing sequence generated by the log.
  2. What if a user wants to edit a comment after it’s already been streamed? – Store edits as new events in the log, update the cache, and push an “edit” message to clients. The DB keeps the latest version.
  3. How do you ensure durability if the cache crashes? – The cache is a hot store; the source of truth is the durable DB. On restart, the read service can rebuild the cache from recent DB rows or replay the log.
  4. How would you add a “reactions” feature? – Treat reactions as separate events on the same Pub/Sub channel. Increment counters in Redis and persist aggregates periodically.
  5. What monitoring would you put in place? – Track write latency, queue depth, cache hit ratio, and consumer lag. Alert on spikes in error rate or sustained high lag.

9. How to Practice This

  1. Sketch the architecture on paper – Start with the API, then add the write log, worker, cache, and read path. Practice explaining each component in under a minute.
  2. Run a mock interview – Use Call Assistant to record yourself answering aloud. Listen back to ensure you keep the flow and tie any anecdotes to your own resume.
  3. Implement a mini‑prototype – Build a simple Node/Go service that writes comments to a Kafka topic and streams them via SSE. Experiment with Redis sorted sets for pagination; this hands‑on experience makes the discussion concrete.

FAQ

  • Q: Do I need a relational database for comments? A: It’s a common choice because it offers strong consistency and flexible queries for pagination and moderation. A NoSQL store can work too, but you’ll need to handle secondary indexes for ordering.

  • Q: How important is exactly‑once delivery? A: For live comments, duplicate delivery is usually tolerable (clients can deduplicate by comment ID). Exactly‑once semantics add complexity and are often unnecessary.

  • Q: Can I skip the cache and read directly from the DB? A: You can, but latency will suffer during spikes. A cache dramatically reduces read latency for hot streams and offloads the DB.

  • Q: What if the moderation service is slow? A: Run it asynchronously. Accept the comment, show it to the user, and hide it later if the filter flags it. This keeps the write path fast while still protecting the community.

Frequently asked questions

Do I need a relational database for comments?

It’s a common choice because it offers strong consistency and flexible queries for pagination and moderation. A NoSQL store can work too, but you’ll need to handle secondary indexes for ordering.

How important is exactly-once delivery?

For live comments, duplicate delivery is usually tolerable (clients can deduplicate by comment ID). Exactly-once semantics add complexity and are often unnecessary.

Can I skip the cache and read directly from the DB?

You can, but latency will suffer during spikes. A cache dramatically reduces read latency for hot streams and offloads the DB.

What if the moderation service is slow?

Run it asynchronously. Accept the comment, show it to the user, and hide it later if the filter flags it. This keeps the write path fast while still protecting the community.

#system design#live comments#architecture#scalability#interview prep#a live comments system