When interviewers ask you to design a file‑storage service like Dropbox, they are looking for three things: a clear set of requirements, a sensible component diagram, and a deep dive into the parts that usually break at scale. Below is a practical walkthrough that you can use in a live interview. The focus is on concrete decisions you can justify, not on vague buzzwords.

1. Clarify Requirements

Start by asking a few clarification questions. This shows you care about scope and helps you avoid building a solution that is too big or too small.

  • Functional
    • Users can upload, download, rename, move, and delete files.
    • Clients sync a local folder with the cloud in near‑real time.
    • Version history is kept for a configurable retention period.
    • Sharing: generate a link or invite collaborators with read/write rights.
  • Non‑functional
    • Scalability – support millions of users and petabytes of data.
    • Latency – upload/download latency should be under a second for typical files (<10 MB).
    • Durability – data loss probability < 10⁻⁹ per year (typical "eleven nines" target).
    • Consistency – eventual consistency for metadata is acceptable, but conflict resolution must be deterministic.
    • Security – data at rest encrypted, access controlled per user.

If the interviewer wants a narrower scope (e.g., only the sync protocol), adjust accordingly.

2. Core Entities and Data Model

EntityPrimary KeyKey Attributes
Useruser_idemail, hashed password, quota, list of device IDs
Filefile_id (content‑addressable hash)owner_id, path, size, MIME type, timestamps, version pointer
Folderfolder_id (also a path)owner_id, parent_folder_id, list of child IDs
Versionversion_idfile_id, version_number, delta (optional), created_at
ShareLinklink_idtarget_file_id, creator_id, permissions, expiry

The most important design decision is to make the file content immutable and identified by a hash of its bytes (e.g., SHA‑256). This enables deduplication and simplifies garbage collection.

3. High‑Level Architecture

+-------------------+          +-------------------+          +-------------------+
|   Client Device   |  Sync →  |   API Gateway     |  RPC →   |   Metadata Service |
+-------------------+          +-------------------+          +-------------------+
          |                               |                               |
          |                               |                               |
          v                               v                               v
+-------------------+          +-------------------+          +-------------------+
|   Sync Engine    |  ↔  Blob Store (object storage)  ↔  |   Auth Service   |
+-------------------+          +-------------------+          +-------------------+
  • API Gateway – entry point, terminates TLS, rate‑limits, routes to micro‑services.
  • Auth Service – issues short‑lived JWTs, validates scopes.
  • Metadata Service – stores file/folder trees, version pointers, share links. A distributed key‑value store (e.g., CockroachDB) works well for strong reads on a per‑user basis.
  • Blob Store – object storage (e.g., S3‑compatible) that holds the immutable file chunks. Use multipart upload for large files.
  • Sync Engine – runs on the client, watches the local folder, generates change events, and talks to the gateway via a bidirectional streaming RPC (gRPC or WebSocket). It also handles conflict resolution using CRDT‑style rules.

4. API Sketch

POST   /files            # multipart upload, returns file_id
GET    /files/{id}       # download
PATCH  /files/{id}       # rename/move, body: {"path": "/new/path"}
DELETE /files/{id}
GET    /folders/{id}/children
POST   /share/{id}       # create share link, returns link URL
GET    /share/{link_id}  # resolve link to file metadata

Keep the API surface small; most operations are CRUD on the metadata service. The heavy lifting (chunking, deduplication) happens behind the scenes.

5. Deep Dive: Hard Parts

5.1. Synchronization & Conflict Resolution

When two devices edit the same file concurrently, you need a deterministic merge strategy. A practical approach is:

  1. Chunk the file into fixed‑size blocks (e.g., 4 MiB).
  2. Compute block hashes; upload only changed blocks.
  3. Version vector per file – each client increments its own counter.
  4. On merge, if version vectors are concurrent, treat the file as conflicted and create a copy with a suffix like "(conflict from Device‑X)".

This mirrors what Dropbox does and is easy to explain.

5.2. Deduplication & Garbage Collection

Because files are content‑addressed, identical blocks from different users point to the same object in the blob store. To reclaim space:

  • Track reference counts per block.
  • Run a periodic job that deletes blocks whose count drops to zero.
  • Store counts in a lightweight KV store; updates are cheap because they are idempotent.

5.3. Scaling Metadata Service

Metadata reads are typically small (a folder listing) but can be hot for popular users. Strategies:

  • Sharding by user_id – isolates workload per tenant.
  • Read‑through cache – in‑memory LRU per API node for recent folder trees.
  • Write‑ahead log – ensures durability without slowing reads.

5.4. Security & Encryption

  • Encrypt each block with a per‑user key derived from their password (or a key stored in a KMS). The gateway never sees plaintext.
  • Share links contain a signed token that encodes permissions; the server validates the signature without looking up a DB row.

6. Trade‑offs & Alternatives

ConcernOption A (Simple)Option B (More Complex)
Metadata StoreSingle relational DBDistributed NewSQL with geo‑replication
Sync ProtocolPolling every few secondsPersistent bidirectional stream (WebSocket)
Conflict HandlingLast‑writer‑winsCRDT‑based merge, user‑editable conflict resolution
StorageSingle region object storeMulti‑region active‑active replication
  • Why choose A? Faster to prototype, easier to reason about consistency.
  • Why choose B? Needed when latency across continents must stay low or when regulatory data‑locality rules apply.

7. Follow‑up Questions Interviewers Often Ask

  1. How would you handle large files (hundreds of MB to GB)?
    • Use multipart upload, resumable chunks, and parallelism. Store chunk hashes in metadata to allow deduplication across large files.
  2. What if a user wants to share a folder with edit rights?
    • Extend the ShareLink model to include a folder ID and a permission enum. The sync engine checks permissions before applying remote changes.
  3. How do you guarantee durability across data‑center failures?
    • Replicate each block to at least three distinct availability zones; use erasure coding for cost efficiency.
  4. Can you make the system eventually consistent for metadata?
    • Yes, but you must expose version vectors to the client so it can detect and resolve conflicts.
  5. How would you test the sync logic?
    • Simulate concurrent edits with a harness that injects network partitions and latency; verify that the final state matches the deterministic merge rules.

8. Putting It All Together

When you present the design, walk the interviewer through:

  1. Requirements – brief recap of functional/non‑functional.
  2. Data model – show the table of core entities.
  3. Component diagram – describe each block’s responsibility.
  4. API – highlight the most used endpoints.
  5. Hard parts – pick one (e.g., conflict resolution) and dive deeper.
  6. Trade‑offs – acknowledge alternatives and why you chose the current design.
  7. Potential extensions – hint at versioning, collaborative editing, or mobile‑first optimizations.

Keep the narrative under 10 minutes; each section should be a sentence or two, with a quick sketch on the whiteboard.

How to practice this

  1. Write the answer aloud – use Call Assistant to capture your spoken flow and keep it on topic.
  2. Mock interview – pair with a peer, swap roles, and enforce a 45‑second limit per section.
  3. Iterate on the diagram – redraw the component diagram from memory until you can do it without looking at notes.

FAQ

  • What is the simplest way to store file contents? Use an object store where each immutable block is addressed by its hash. This enables deduplication and easy garbage collection.
  • Do I need a separate service for version history? Not necessarily; you can store version pointers in the metadata service and keep deltas as optional blocks.
  • How does Dropbox handle offline edits? The client queues changes locally, applies them when connectivity returns, and resolves conflicts using version vectors.
  • Is eventual consistency acceptable for a file‑storage service? Yes, as long as the client can detect divergent versions and present a deterministic conflict resolution to the user.

Frequently asked questions

What is the simplest way to store file contents?

Use an object store where each immutable block is addressed by its hash. This enables deduplication and easy garbage collection.

Do I need a separate service for version history?

Not necessarily; version pointers can live in the metadata service, with optional delta blocks stored alongside the main blobs.

How does Dropbox handle offline edits?

The client queues changes locally, syncs them when online, and resolves conflicts with version vectors, creating conflict copies when needed.

Is eventual consistency acceptable for a file‑storage service?

Yes, as long as the client can detect divergent versions and apply a deterministic conflict‑resolution strategy.

#system design#file storage#dropbox#sync#architecture#a file storage service like Dropbox