Coinbase’s system design interview is a deep dive into how you would build services that handle money, regulatory pressure, and massive traffic spikes. The round lasts about 45‑60 minutes, and you’ll be expected to sketch architecture, discuss component choices, and walk through failure scenarios. Below is a practical map of what the interview looks like, how interviewers score you, and how to prepare without drowning in hype.
What the Round Usually Covers
| Area | Typical focus | Why it matters to Coinbase |
|---|---|---|
| Scalability | Horizontal scaling, sharding, load balancing | Crypto trading volumes can jump 10x during market events. |
| Security & Compliance | Encryption, auth, audit logs, KYC data handling | Regulatory bodies demand strong controls over user funds. |
| Data Consistency | CAP trade‑offs, eventual vs strong consistency | Wallet balances must be accurate across devices. |
| Observability | Metrics, alerting, tracing | Outages cost money and reputation; rapid detection is crucial. |
| Operational Simplicity | Deployability, rollback strategy | Engineers need to push fixes quickly under pressure. |
Interviewers typically start with a high‑level problem statement, ask you to clarify requirements, then move through the classic design stages: requirements, high‑level components, detailed flow, bottlenecks, and failure handling. They will probe for trade‑offs and expect you to justify decisions with concrete reasoning rather than vague preferences.
The Rubric Interviewers Use
- Understanding Requirements – Do you ask clarifying questions about latency, durability, and compliance? Do you identify hidden constraints like regulatory auditability?
- System Decomposition – Are the major services (e.g., API gateway, transaction service, ledger, notification) clearly separated with well‑defined interfaces?
- Data Modeling & Consistency – Do you choose the right consistency model for balances vs market data? Can you explain why a ledger might be append‑only?
- Scalability & Performance – Do you discuss sharding keys, caching layers, and async pipelines? Are you aware of typical traffic patterns for crypto exchanges?
- Security & Reliability – Do you cover encryption at rest, TLS, rate limiting, and replay‑attack protection? Do you describe fallback paths for partial outages?
- Observability & Ops – Do you mention metrics, distributed tracing, and automated rollbacks? Can you sketch a simple alerting rule for transaction failures?
- Communication – Do you keep the conversation structured, summarize frequently, and stay on topic?
Each dimension is scored roughly on a low‑medium‑high scale. The overall decision hinges on whether you demonstrate a balanced view: you can scale, you can secure, and you can operate the system.
Example Prompt #1: Design a Crypto Wallet Service
Prompt (paraphrased) – "Design a service that lets users create, store, and transfer multiple cryptocurrencies. Assume the system must support high‑frequency small transfers and comply with AML/KYC regulations."
High‑Level Sketch
- API Gateway – Authenticated entry point, throttles per‑user rate.
- Wallet Service – Handles creation, balance queries, and address generation. Stores metadata in a relational DB for quick look‑ups.
- Ledger Service – Append‑only, immutable log of all transfers. Uses a distributed log (e.g., Kafka) for durability and replay.
- Balance Service – Materialized view built from the ledger, refreshed via stream processing (Flink/Beam). Provides eventual consistency for read‑heavy workloads.
- Compliance Service – Hooks into transaction flow to run AML checks; stores audit records in a tamper‑evident store.
- Notification Service – Pushes email/SMS/webhook alerts on transaction events.
Key Trade‑offs
- Consistency vs Latency – For small transfers, users expect sub‑second confirmation. You can show the balance optimistically, then reconcile with the ledger asynchronously.
- Sharding Strategy – Partition by user ID to keep hot wallets isolated, reducing cross‑shard traffic.
- Security – Private keys never touch the backend; they are generated client‑side and stored encrypted in the user's device.
Failure Scenarios
- Ledger outage – Buffer incoming transactions in a local queue, replay once the log recovers.
- Compliance service latency – Use a fast‑path that allows low‑value transfers to bypass full AML checks, with a post‑process audit.
Sample Answer (45‑90 s)
"I’d start by separating the user‑facing API from the core ledger. The API gateway would enforce authentication and rate limits. The wallet service would create accounts and expose balance queries, storing static metadata in a relational store for low‑latency reads. All transfers would be written to an append‑only ledger backed by a distributed log; this gives us an immutable audit trail required for regulators. A separate balance service would materialize the ledger into per‑user balances using a streaming processor, providing near‑real‑time reads while tolerating brief inconsistencies. AML checks would run as a side‑car service that tags each transaction; low‑value transfers could be fast‑tracked with a later batch audit. In case the ledger goes down, we’d buffer new events locally and replay them once the log is healthy. Observability would be built in via metrics on queue depth, latency, and error rates, with alerts wired to a run‑book for immediate rollback."
Example Prompt #2: Design a Real‑Time Crypto Price Feed
Prompt (paraphrased) – "Build a system that streams live price updates for dozens of cryptocurrencies to millions of subscribers with low latency and high reliability."
High‑Level Sketch
- Market Data Ingestion – Connectors to exchanges via WebSocket; normalize data into a common schema.
- Aggregator Service – Merges order‑book snapshots, calculates mid‑price, and applies smoothing algorithms.
- Publish‑Subscribe Layer – Uses a high‑throughput pub/sub system (e.g., NATS or Pulsar) to fan‑out updates.
- Cache Layer – Edge caches (CDN or regional Redis) store the latest price per symbol for sub‑millisecond reads.
- Client API – WebSocket or gRPC endpoint that streams updates; supports subscription filters.
Key Trade‑offs
- Latency vs Throughput – Direct push from aggregator to clients yields the lowest latency but can overload the network; a tiered pub/sub with edge caching balances both.
- Data Freshness – Prices are volatile; you can tolerate a few milliseconds of staleness if it reduces load.
- Reliability – Use replay‑able logs for each symbol so a new subscriber can catch up without missing a beat.
Failure Scenarios
- Exchange disconnect – Fallback to the last known price and flag the symbol as stale until reconnection.
- Pub/Sub node failure – Deploy multiple brokers in a cluster; clients reconnect automatically.
Sample Answer (45‑90 s)
"I’d architect the feed around a pipeline that ingests market data from exchange APIs, normalizes it, and feeds it into an aggregator that computes the latest price per coin. The aggregator would publish updates to a high‑throughput pub/sub system, which fans out to regional edge caches. Clients would connect via WebSocket and receive a stream filtered to the symbols they care about. To keep latency low, the edge cache would serve the most recent price directly, while a replay log ensures new subscribers can backfill missed updates. If an exchange drops, we’d continue broadcasting the last known price and mark the feed as stale until reconnection. Observability would include per‑symbol lag metrics and broker health checks, with alerts that trigger automated failover to a secondary data source."
How to Practice This
- Weekly Design Sprint – Pick a new prompt each week, sketch the architecture on a whiteboard, and record a 5‑minute explanation. Review the recording for clarity and completeness.
- Rubric Drill – After each mock, score yourself against the seven rubric items. Identify the lowest‑scoring area and focus the next sprint on that dimension.
- Mock Interview with Call Assistant – Use Call Assistant to rehearse your answer aloud. It will capture the flow, keep follow‑up questions on topic, and surface gaps where your story diverges from your résumé.
FAQ
What level of detail does Coinbase expect for data stores? They look for a clear justification of choice—relational for metadata, append‑only logs for immutable transaction history, and eventual‑consistent caches for read‑heavy paths.
How important is regulatory knowledge in the design round? Very important. You should mention audit logs, immutable storage, and how data residency requirements could affect component placement.
Can I bring diagrams to the interview? Yes, but they should be drawn in real time on a shared whiteboard tool. Pre‑made slides are usually discouraged because they hide your thought process.
What is a good fallback strategy for a critical service like the ledger? Buffer incoming writes locally, use a replicated log for durability, and have an automated replay mechanism that restores consistency once the primary path recovers.
Frequently asked questions
What level of detail does Coinbase expect for data stores?
They look for a clear justification of choice—relational for metadata, append‑only logs for immutable transaction history, and eventual‑consistent caches for read‑heavy paths.
How important is regulatory knowledge in the design round?
Very important. Mention audit logs, immutable storage, and data‑residency considerations that affect where components are deployed.
Can I bring diagrams to the interview?
Yes, but they should be drawn live on a shared whiteboard. Pre‑made slides can hide your reasoning and are usually discouraged.
What is a good fallback strategy for a critical service like the ledger?
Buffer writes locally, use a replicated log for durability, and have an automated replay process that restores consistency once the primary path recovers.
#Coinbase#system design#interview prep#architecture#crypto