When an interview asks you to design a photo‑sharing service like Instagram, the conversation usually starts with a high‑level vision and quickly drills into the parts that stress scalability, latency, and consistency. Below is a practical walkthrough you can follow in real time, with concrete examples you can adapt to your own background.

1. Clarify Scope and Requirements

Functional requirements

  • Users can upload photos with optional captions.
  • Users can follow other users and see a chronological or algorithmic feed.
  • Users can like and comment on photos.
  • Basic profile management (username, avatar, bio).
  • Search for users or hashtags.

Non‑functional requirements

  • Availability: High‑availability for read‑heavy workloads; tolerate occasional write latency.
  • Latency: Feed reads should return within a few hundred milliseconds for a smooth UI.
  • Scalability: System must handle growth from thousands to millions of daily active users.
  • Consistency: Likes and comments need strong consistency for a given photo; eventual consistency is acceptable for the feed.
  • Security & privacy: Authenticated access, private accounts, and content moderation.

Ask the interviewer which of these are most important for the scenario. If they say “focus on the feed generation,” you can de‑emphasize profile editing details.

2. Identify Core Entities and Data Model

EntityKey FieldsTypical Access Patterns
Useruser_id, username, avatar_url, bio, follower_set, following_setRead‑heavy (profile view), write‑heavy for follow/unfollow
Photophoto_id, owner_id, url, caption, timestamp, like_count, comment_countWrite on upload, read for feed, incremental updates for likes/comments
Likeuser_id, photo_id, timestampWrite‑heavy, read for per‑photo like list
Commentcomment_id, photo_id, user_id, text, timestampWrite‑heavy, read for comment thread
Feeduser_id, list<photo_id> (ordered)Read‑heavy, occasional rebuild on follow change

The Feed entity is a derived view; you can store it as a materialized list per user or generate it on‑the‑fly.

3. High‑Level Architecture

+-------------------+      +-------------------+      +-------------------+
|   API Gateway     | ---> |   Auth Service    | ---> |   User Service    |
+-------------------+      +-------------------+      +-------------------+
        |                         |                         |
        v                         v                         v
+-------------------+      +-------------------+      +-------------------+
|   Photo Service   | ---> |   Media Storage   |      |   Feed Service    |
+-------------------+      +-------------------+      +-------------------+
        |                         |                         |
        v                         v                         v
+-------------------+      +-------------------+      +-------------------+
|   Like Service    | ---> |   Notification    |      |   Search Service  |
+-------------------+      +-------------------+      +-------------------+
  • API Gateway routes requests, enforces rate limits, and logs metrics.
  • Auth Service validates JWTs or OAuth tokens.
  • User Service stores profile data in a relational store for strong consistency.
  • Photo Service handles upload metadata, writes to a durable object store (e.g., S3‑compatible), and queues a background job for thumbnail generation.
  • Media Storage is a CDN‑backed object store; reads are served directly from edge locations.
  • Feed Service maintains per‑user timelines. Two common strategies are discussed later.
  • Like Service updates counters atomically and pushes events to a message bus.
  • Notification Service consumes events to alert followers.
  • Search Service indexes hashtags and usernames using a search engine like Elasticsearch.

4. Deep Dive: Photo Upload Pipeline

  1. Client → API Gateway: multipart POST with auth token.
  2. Gateway → Photo Service: validates size, content type, and extracts metadata.
  3. Photo Service → Object Store: writes the original image to a bucket; returns a storage URL.
  4. Background Workers: read from a queue, generate thumbnails (different sizes), and store them alongside the original.
  5. Metadata DB: a NoSQL table (e.g., DynamoDB) stores photo_id, owner_id, url, timestamp, and counts.

Why a queue? It decouples the upload latency from image processing, keeping the user‑facing request fast.

5. Hard Problem #1 – Feed Generation at Scale

Pull‑Based (On‑Demand) Model

  • When a user opens the app, the Feed Service queries recent photos from the users they follow, merges them by timestamp, and returns the top N.
  • Pros: Minimal storage overhead; easy to add new follow relationships.
  • Cons: High read latency; each request hits many user photo tables.

Push‑Based (Fan‑out) Model

  • On photo upload, the system pushes the photo_id to the feed list of each follower (fan‑out write).
  • Pros: Reads are cheap – just fetch a pre‑computed list.
  • Cons: Write amplification when a user has many followers; may require sharding and background workers.

Hybrid Approach

  • Use push for users with a moderate follower count (e.g., <10k) and pull for “celebrity” accounts with massive followings.
  • Store a fan‑out queue for high‑fan‑out accounts; generate a global recent‑photos stream that pull‑based readers can merge.

Data Store Choices

  • Redis Sorted Sets for per‑user timelines (fast range queries, TTL support).
  • Cassandra or ScyllaDB for wide rows when follower count is large.
  • Kafka streams can replay events to rebuild feeds if needed.

6. Hard Problem #2 – Media Delivery and CDN Caching

  • Store original images in a durable object store; generate multiple resolutions (e.g., 1080p, 720p, thumbnail).
  • Use signed URLs for private content; public images can be cached aggressively at edge locations.
  • Cache‑control headers dictate how long browsers keep images; set short TTL for stories that expire quickly.
  • For “stories” (ephemeral content), purge CDN objects after 24 hours using an automated invalidation job.

7. Trade‑offs and Common Follow‑Up Questions

TopicTypical Interview Follow‑UpSample Answer Angle
Consistency"How do you ensure like counts are accurate?"Use atomic increments in the NoSQL store; optionally write to a write‑ahead log and reconcile with background compaction.
Scaling"What happens if a user with millions of followers uploads a photo?"Switch to pull‑based retrieval for that user, store the photo in a global recent‑photos stream, and let readers merge with their own follow list.
Data Privacy"How do you support private accounts?"Store a visibility flag per user; feed queries filter out private accounts unless the requester follows them.
Rate Limiting"How do you prevent abuse of the upload API?"Enforce per‑user token bucket limits at the API gateway; throttle heavy uploaders.
Testing"How would you test the feed generation logic?"Use property‑based tests that simulate follow graphs and verify that the top‑N returned photos respect timestamps and visibility rules.

When the interviewer asks about a specific component, walk through the data flow, mention the technology choices, and discuss the trade‑offs you considered.

8. Sample Answer Template (45‑90 seconds)

"Sure, let me walk through a high‑level design. The system has an API gateway that routes authenticated requests to dedicated services: a Photo Service for uploads, a Feed Service for timeline generation, and a Like Service for interactions. Photos are stored in an object store behind a CDN, and metadata lives in a NoSQL table. For the feed, we use a hybrid push‑pull model: most users get a pre‑computed timeline stored in Redis sorted sets, while accounts with massive followings fall back to on‑demand merging of recent posts. This keeps reads fast for the majority of cases and avoids write amplification for celebrity accounts. Likes are stored with atomic counters, and a message bus notifies followers in near real‑time. The whole pipeline is decoupled with queues, so the user sees the uploaded photo instantly while thumbnails are generated asynchronously."

9. How to Practice This

  1. Sketch the diagram on paper before you speak. Focus on the request flow for one core operation (e.g., photo upload).
  2. Record yourself answering a typical prompt and listen back. Use Call Assistant to keep the answer grounded in a real project you’ve worked on.
  3. Iterate on trade‑offs: pick a component, ask “what if we change this technology?” and articulate the impact on latency, cost, and complexity.

FAQ

  • What is the simplest way to store per‑user feeds? Use a Redis sorted set keyed by feed:{user_id} where the score is the photo timestamp. It supports range queries and expiration, making it ideal for a read‑heavy timeline.

  • How do you handle photo deletion? Mark the photo as deleted in the metadata store, remove the entry from followers' feeds (or let a background job clean it), and delete the object from storage after a grace period.

  • Why not store all photos in a relational database? Relational tables struggle with high write throughput and large binary blobs. Separating metadata (SQL/NoSQL) from binary data (object store) gives better scalability and cheaper storage.

  • Can the same design support video? Yes, replace the image processing pipeline with a transcoding pipeline that produces multiple bitrate renditions, and store video metadata similarly. The feed logic remains unchanged.

Frequently asked questions

What is the most important metric to monitor for a photo‑sharing service?

Read latency for the feed is usually the primary KPI because it directly affects user experience. You also keep an eye on upload success rate and storage utilization.

How do you ensure the system stays available during a data‑center outage?

Deploy services across multiple availability zones, use a replicated object store, and place a read‑only replica of the feed in each zone. The API gateway can route traffic to the healthy zone automatically.

What security measures protect user‑generated content?

Authenticate every request with JWTs, enforce ACLs on private accounts, scan uploaded media for malware, and apply rate limits to prevent abuse.

When should you switch from push‑based to pull‑based feed generation?

When a user’s follower count exceeds the threshold where fan‑out writes become costly (often in the tens of thousands). At that point, generating the feed on demand avoids massive write amplification.

#system design#photo sharing#instagram#scalability#architecture#a photo sharing service like Instagram