When you sit down for a system design interview, the first thing the interviewer expects is a clear, shared understanding of what the service does. For a YouTube‑style video streaming platform, that means defining the user‑facing features, the performance goals, and the constraints that shape the architecture. Below is a practical walk‑through you can use in a real interview. It follows the flow most interviewers use: requirements gathering → high‑level sketch → deep dive on the toughest pieces → trade‑off discussion → possible follow‑ups.
1. Clarify Functional Requirements
Start by asking the interviewer to confirm the scope. Typical functional points for a video streaming service include:
- Upload: Authenticated users can upload a video file of arbitrary length.
- Processing: The system generates multiple bitrate renditions and extracts thumbnails.
- Metadata: Users can set title, description, tags, privacy settings, and can edit them later.
- Search & Discovery: Browse by channel, tag, or recommendation feed.
- Playback: Viewers request a video, receive a manifest, and stream the appropriate bitrate.
- Comments / Likes: Basic interaction primitives.
- Monetization: Optional ad insertion points (can be omitted for the core design).
If the interview is time‑boxed, you can safely say you will focus on upload, processing, storage, CDN delivery, and basic search.
2. Clarify Non‑Functional Requirements (NFRs)
NFRs drive most of the architectural decisions. Ask the interviewer for typical values, or use placeholders:
| Requirement | Typical Target |
|---|---|
| Availability | 99.9% for video playback, higher for upload pipeline |
| Latency | < 200 ms for manifest fetch, < 2 s for initial video chunk |
| Throughput | Handle a steady stream of uploads (denoted R requests/second) and millions of concurrent viewers |
| Consistency | Strong consistency for metadata edits; eventual consistency for view counts |
| Scalability | Horizontal scaling for both upload and playback paths |
| Security | Authenticated upload, public read‑only playback |
You don’t need exact numbers; the placeholder R lets you talk about scaling without inventing traffic.
3. Identify Core Entities and Data Model
A minimal data model helps you reason about storage and APIs.
classDiagram
class User {
+userId: UUID
+email: string
+profileInfo
}
class Video {
+videoId: UUID
+ownerId: UUID
+title: string
+description: string
+tags: []string
+privacy: enum
+uploadTime: timestamp
}
class VideoFile {
+videoId: UUID
+resolution: enum
+format: string
+url: string
}
class Comment {
+commentId: UUID
+videoId: UUID
+authorId: UUID
+content: string
+timestamp
}
User "1" --> "*" Video : owns
Video "1" --> "*" VideoFile : renditions
Video "1" --> "*" Comment : has
- Video stores metadata.
- VideoFile records each transcoded rendition (e.g., 1080p MP4, 720p WebM).
- Comment is a separate entity to keep write traffic independent of playback.
4. Sketch a High‑Level Architecture
A textual diagram is easy to draw on a whiteboard and covers the major services.
Client <---> Load Balancer <---> API Gateway <---> Auth Service
|
+-- Upload Service <---> Object Store (raw uploads)
|
+-- Transcoding Service <---> Worker Queue <---> Compute Cluster
|
+-- Metadata Service (SQL/NoSQL) <---> Search Index
|
+-- CDN Edge Nodes <---> Object Store (processed files)
Key components:
- Load Balancer distributes inbound traffic.
- API Gateway routes requests to micro‑services.
- Auth Service validates JWTs or OAuth tokens.
- Upload Service receives multipart uploads, stores the original file in an object store (e.g., S3‑compatible), and enqueues a job.
- Transcoding Service pulls jobs from the queue, runs codecs on a compute cluster, and writes renditions back to the object store.
- Metadata Service holds searchable fields; a secondary search index (e.g., Elasticsearch) powers keyword search.
- CDN Edge Nodes cache the final video files for low‑latency playback.
5. Deep Dive: The Hard Parts
5.1 Video Ingestion & Transcoding
Problem: A single upload can be several gigabytes and must be split into multiple bitrate renditions.
Solution pattern:
- Chunked upload – client sends the file in small parts; the service assembles them in the object store.
- Job creation – once the upload completes, a message (e.g., to a Pub/Sub system) triggers a transcoding job.
- Worker pool – each worker pulls the raw file, runs a pipeline (e.g., FFmpeg) to produce HLS/DASH segments.
- Storage layout – store renditions under a predictable key:
videos/{videoId}/1080p/segment0.ts.
Trade‑offs:
- Cold start latency vs. cost: Keep a pool of warm containers to reduce job start time, at the expense of higher idle cost.
- Quality vs. speed: Use two‑pass encoding for best quality, but a single pass is faster for low‑priority uploads.
5.2 Content Delivery Network (CDN) Integration
The CDN is the bridge between the object store and the viewer. In an interview you can discuss two common models:
| Model | Pros | Cons |
|---|---|---|
| Push‑based (pre‑populate edge caches) | Guarantees low latency for hot videos; simple cache‑control logic. | Wastes bandwidth on rarely‑watched content. |
| Pull‑based (edge fetches on‑demand) | Efficient for long‑tail content; only populates what viewers request. | First‑request latency can be higher; requires robust origin. |
Most large platforms use a hybrid: hot videos are pre‑cached, the rest rely on pull.
5.3 Search & Recommendation Index
Metadata lives in a relational store for ACID guarantees, but search needs full‑text capabilities. The usual pattern is:
- Write‑through: On every metadata change, emit an event to a stream.
- Indexer: Consume the stream and update a search engine (e.g., OpenSearch).
- Query path: API service forwards search queries to the index, merges results with permission checks.
Hard part: Keeping the index eventually consistent without exposing stale results. You can mention version stamps on documents and a fallback to the primary store if the index misses a recent edit.
5.4 Consistency for View Counts
Counters that increment on every playback can become hot spots. Common approaches:
- Sharded counters: Split the counter across many keys; aggregate on read.
- Approximate counting (HyperLogLog) for unique viewers.
- Eventual consistency: Accept that the displayed count may lag behind real traffic.
Explain why exact real‑time counts are rarely required for user experience, and how sharding helps scale.
6. Trade‑Off Discussion
Interviewers love to see you weigh options. Cover at least three dimensions:
- Latency vs. Cost – Keeping a large pool of warm transcoding workers reduces start‑up latency but raises idle cost. A pull‑based CDN reduces storage cost but adds first‑byte latency.
- Strong vs. Eventual Consistency – Metadata edits need strong consistency (user expects immediate change). View counts can be eventually consistent, allowing sharded counters.
- Monolithic vs. Micro‑services – A monolith speeds up initial MVP but makes scaling individual parts (e.g., transcoding) harder. A micro‑service decomposition adds operational overhead but isolates scaling concerns.
7. Typical Follow‑Up Questions
Interviewers often probe deeper after the high‑level sketch. Be ready for:
- How would you handle live streaming? – Explain ingest via RTMP, segmenting into short HLS/DASH chunks, and using a low‑latency CDN.
- How do you protect copyrighted content? – Discuss DRM (e.g., Widevine), tokenized URLs, and watermarking.
- What if a video becomes viral and spikes traffic? – Talk about autoscaling the CDN edge, pre‑warming transcoding workers, and rate‑limiting upload to protect the origin.
- How would you implement personalized recommendations? – Mention a separate ML pipeline that consumes watch events, builds user embeddings, and serves scores via a low‑latency inference service.
- How do you ensure data privacy across regions? – Store user‑generated content in regions that satisfy regulatory constraints; route requests through region‑aware load balancers.
8. Sample Answer Template (45‑90 seconds)
"Sure, let me outline a YouTube‑style service. First, the functional core includes upload, transcoding into multiple bitrate renditions, storing metadata, and delivering video via a CDN. Non‑functional goals are high availability for playback, sub‑second latency for manifest fetch, and horizontal scalability for both upload and view traffic. The architecture splits into an API gateway, an upload service that writes the raw file to object storage, a worker queue that triggers transcoding on a compute cluster, a metadata service backed by a relational store with a secondary search index, and edge‑cached CDN nodes for playback. The hardest pieces are the transcoding pipeline—because we need to spin up workers quickly without over‑provisioning—and the CDN strategy, where we balance push‑based pre‑caching of hot videos against pull‑based delivery for the long tail. Trade‑offs involve latency versus cost, consistency versus scalability, and monolith versus micro‑service design. If the interviewer asks about viral spikes, I’d mention autoscaling the edge and sharding view counters. Does that cover the scope you had in mind?"
9. How to Practice This
- Sketch the flow aloud – Use a tool like Call Assistant to rehearse your answer, keeping each component under a minute.
- Pick a hard sub‑system (e.g., transcoding) and write a short whiteboard diagram, then explain trade‑offs to a friend.
- Simulate follow‑up questions – Have a peer ask you about live streaming, DRM, or scaling; answer each in 30‑second bursts to build agility.
FAQ
Q: Do I need to design the UI as part of the system design? A: Usually not. Focus on backend services, data flow, and scalability. Mention UI only to clarify user actions.
Q: How much detail should I give about the video codec? A: Briefly note that H.264/H.265 are common, and that the transcoding service abstracts the codec choice. Deep codec internals are out of scope.
Q: Should I include ad insertion in the design? A: Only if the interviewer explicitly asks. Otherwise, treat it as an optional plug‑in that reads from a separate ad‑service.
Q: What if the interviewer asks for exact traffic numbers? A: Use placeholders like R for request rate and discuss scaling behavior (e.g., "the system should handle R uploads per second by adding more workers").
Frequently asked questions
What are the minimal functional pieces needed for a video streaming service?
Upload, transcoding into multiple renditions, metadata storage, search, and CDN‑based playback. Optional pieces include comments, likes, and ad insertion.
Why do we use a separate search index instead of the primary database?
Full‑text search and ranking require inverted indexes and scoring algorithms that relational databases don’t provide efficiently. A write‑through sync keeps the index up to date while preserving strong consistency for metadata.
How can view counters be scaled without becoming a hotspot?
Shard the counter across many keys and aggregate on read, or use approximate counting structures. This spreads write load and avoids a single hot partition.
When is a push‑based CDN better than a pull‑based one?
Push‑based caching works well for hot videos that receive a lot of traffic, guaranteeing low latency at the cost of extra bandwidth. Pull‑based caching is more efficient for the long tail where pre‑populating caches would waste resources.
#system design#video streaming#architecture#scalability#interview#a video streaming service like YouTube