DoorDash’s system design interview is a deep dive into how you think about large‑scale, latency‑sensitive services. The interview lasts about 45‑60 minutes and is usually conducted by a senior engineer or a tech lead who has built parts of the core platform. The goal isn’t to produce production‑ready code; it’s to see how you break a vague problem into concrete components, reason about trade‑offs, and keep the conversation anchored to the product’s business goals.
What the round typically covers
| Area | What interviewers probe |
|---|---|
| Scope definition | How you ask clarifying questions, identify user personas, and decide what to include or exclude. |
| Core components | Your ability to sketch services, data stores, queues, and APIs that satisfy the functional requirements. |
| Scalability & performance | Choices around sharding, caching, load balancing, and async processing to handle peak traffic. |
| Reliability | Redundancy, failover, data replication, and monitoring strategies. |
| Data modeling | How you store and retrieve the key entities (orders, drivers, restaurants) while supporting the required queries. |
| Consistency vs. latency | When you favor eventual consistency versus strong consistency, and why. |
| Product sense | Connecting technical decisions back to DoorDash’s metrics like delivery time, churn, and merchant revenue. |
Interviewers often follow a loosely shared rubric that emphasizes:
- Clarity of requirements – Did you surface the most important functional and non‑functional constraints?
- System decomposition – Are the major services well separated and do they have clear responsibilities?
- Trade‑off justification – Do you explain why you chose a particular database, queue, or replication factor?
- Bottleneck identification – Can you spot the part of the system that would limit throughput and propose mitigations?
- Communication – Do you keep the discussion structured and respond to follow‑up prompts without veering off topic?
Typical prompt #1: Real‑time order tracking
Prompt (paraphrased): Design a system that lets customers see the live location of their driver and estimated time of arrival, updating every few seconds.
High‑level walk‑through
- Clarify requirements
- Frequency of location updates (e.g., every 5 seconds).
- Scale: peak concurrent orders (tens of thousands) and geographic coverage.
- Latency target: end‑to‑end < 2 seconds for the UI.
- Fault tolerance: what happens if a driver’s device drops connectivity?
- Key components
- Driver SDK that streams GPS coordinates to a Location Service via a lightweight protocol (e.g., gRPC over cellular).
- Location Service writes points to a high‑throughput time‑series store (e.g., a partitioned Kafka topic or a purpose‑built geo‑store).
- Realtime API reads the latest point per driver from an in‑memory cache (Redis or a custom sharded cache) and pushes updates to the client via WebSocket.
- ETA Service consumes the location stream, applies routing logic (e.g., OSRM or a proprietary routing engine), and writes ETA predictions to the same cache.
- Data flow
- Driver → gRPC → Location Service → Kafka → Consumer (cache updater) → Redis → WebSocket → Mobile UI.
- Scalability tricks
- Partition Kafka by city or driver ID to spread load.
- Use a CDN‑backed WebSocket edge to offload connection handling.
- Batch cache writes per second to reduce write amplification.
- Reliability
- Replicate Kafka topics across three brokers for durability.
- Store the last known location in a durable store (e.g., DynamoDB) for fallback if the cache is lost.
- Heartbeat messages from the driver to detect disconnections; fall back to “last known location” UI.
- Trade‑offs
- Strong consistency vs. latency: we accept eventual consistency for the map view because a few hundred‑millisecond lag is tolerable.
- Geo‑store vs. generic DB: a specialized geo‑index reduces query latency for “nearest driver” calculations.
Sample answer snippet (45‑90 seconds)
“First I’d ask how often the UI needs fresh coordinates and what latency we’re targeting. Assuming a 5‑second interval and sub‑2‑second end‑to‑end latency, I’d have the driver SDK push GPS points over gRPC to a stateless Location Service that writes to a partitioned Kafka topic. A consumer would update an in‑memory cache keyed by driver ID, and a WebSocket layer would push the latest point to the customer app. For ETA we’d feed the same stream into a routing service that outputs predictions back into the cache. To keep the system reliable, Kafka would be replicated across three brokers and we’d persist the last known point in a durable store for fallback. The main trade‑off is choosing eventual consistency for the map view, which lets us keep latency low while still delivering a smooth experience.”
Typical prompt #2: Restaurant onboarding pipeline
Prompt (paraphrased): Design a backend service that lets new restaurants sign up, upload menus, and go live on DoorDash within a day.
High‑level walk‑through
- Clarify requirements
- Expected onboarding time (e.g., 24 hours max).
- Volume: spikes of a few hundred new restaurants per day during promotions.
- Data validation: menu items, pricing, dietary tags.
- Integration points: payment processor, fraud detection, and internal catalog service.
- Core services
- Restaurant API – REST endpoints for sign‑up, menu upload, and status checks.
- Validation Service – Synchronous checks for schema compliance; async enrichment (e.g., image processing) via a task queue.
- Catalog Service – Persists approved menu data into a relational store (e.g., PostgreSQL) and publishes events to a downstream catalog cache.
- Workflow Orchestrator – Coordinates steps using a state machine (e.g., AWS Step Functions or an open‑source equivalent).
- Data flow
- Restaurant → HTTPS POST → Restaurant API → Validation Service → (if async) SQS → Worker → Catalog Service → Event Bus → Search Index.
- Scalability
- Autoscale API front‑ends behind an ALB.
- Partition the validation queue by restaurant region to isolate spikes.
- Use read‑replicas for catalog queries that power the UI.
- Reliability
- Idempotent writes to the catalog DB to survive retries.
- Dead‑letter queues for failed validation jobs.
- Monitoring of step durations; alert if any step exceeds a threshold.
- Trade‑offs
- Sync vs. async validation: simple schema checks can be sync; heavy image processing is async to keep the API responsive.
- Relational vs. NoSQL: menu data benefits from relational constraints (foreign keys, transactions), so a SQL store is preferred despite the slight scaling overhead.
Sample answer snippet (45‑90 seconds)
“I’d start by confirming the SLA: we need a restaurant live within 24 hours, even during promotional spikes. The flow would be a RESTful sign‑up API that writes the basic profile to a relational store, then hands off the menu upload to a validation service. Simple schema checks happen synchronously; heavy tasks like image resizing go onto a message queue. A workflow orchestrator tracks each step, persisting state so we can resume after failures. Once validated, the menu is written to the catalog DB and an event is emitted to update the search index. To keep the system resilient we’d make all writes idempotent, use dead‑letter queues for problematic jobs, and autoscale the API tier behind a load balancer. The main trade‑off is using a relational DB for strong consistency at the cost of a bit more scaling effort, which we accept because menu integrity is critical for the customer experience.”
How to practice this
- Sketch on a whiteboard – Pick a recent DoorDash‑related feature (e.g., “instant delivery”) and draw the end‑to‑end flow. Focus on components, data stores, and async boundaries.
- Iterate with a peer – Run a mock interview where the partner plays the interviewer. Use the same cadence as a real call: ask clarifying questions, present a high‑level diagram, and drill into trade‑offs.
- Rehearse with Call Assistant – Record yourself answering a prompt aloud, let the assistant surface follow‑up questions, and refine your story so it stays grounded in your resume and stays on topic.
FAQ
- What level of detail is expected for data stores? You should name the type of store (e.g., “partitioned Kafka for high‑throughput streams” or “PostgreSQL for relational consistency”) and explain why it fits the access pattern, but you don’t need to enumerate exact schema fields.
- How much code should I write? Minimal pseudo‑code is fine if it clarifies an algorithm (e.g., a simple hash‑based sharding function). The focus should remain on architecture, not implementation.
- Do I need to cover security? Mention authentication (OAuth) and data encryption at rest and in transit if the prompt involves sensitive data, but a deep dive isn’t required unless the interviewer asks.
- What if I get stuck on a component? Pause, restate the requirement that component satisfies, and propose a high‑level alternative. Interviewers value honest reasoning over guessing.
Frequently asked questions
What level of detail is expected for data stores?
Name the store type (e.g., Kafka, Redis, PostgreSQL) and justify its fit for the access pattern. Detailed schema isn’t required unless the interviewer asks.
How much code should I write?
Only minimal pseudo‑code to illustrate an algorithm or sharding logic. The interview emphasizes architecture, not full implementation.
Do I need to cover security?
Briefly note authentication and encryption if the problem involves sensitive data. Dive deeper only if the interviewer probes.
What if I get stuck on a component?
Pause, restate the requirement that component addresses, and suggest a high‑level alternative. Transparency in reasoning is preferred over guessing.
#DoorDash#system design#interview prep#architecture#scalability