You’re asked to design a collaborative document editor that feels like Google Docs. The interview will swing between high‑level product thinking and low‑level engineering details, so start by clarifying scope, then walk through the core components, and finally dive into the tricky parts like concurrency control and scaling.
Clarify Requirements First
Functional requirements
- Real‑time editing: multiple users see each other's changes within a second.
- Rich text support: headings, lists, tables, images, and basic formatting.
- Version history: ability to revert to any prior version.
- Access control: owners can invite editors/viewers, revoke access, and set permissions.
- Offline editing: changes sync when the client reconnects.
- Search: find documents by title, content, or tags.
Non‑functional requirements
- Low latency: sub‑second round‑trip for edit propagation.
- High availability: tolerate node failures without losing edits.
- Scalability: support thousands of concurrent editors on popular documents and millions of idle documents.
- Consistency: eventual consistency is acceptable for offline edits, but the UI must appear consistent to active users.
- Security & privacy: encryption at rest and in transit, strict ACL checks.
Tip: If the interviewer asks for a specific SLA, give a reasonable range (e.g., 99.9% uptime, 200 ms 99th‑percentile latency) and note that exact numbers depend on the target market.
Core Entities and Data Model
| Entity | Key fields | Reason it matters |
|---|---|---|
| Document | doc_id, owner_id, title, created_at, updated_at | Primary object, used for ACL checks and indexing |
| Revision | rev_id, doc_id, author_id, ops, timestamp | Enables version history and conflict resolution |
| User | user_id, email, display_name, profile | Authentication and permission mapping |
| Permission | doc_id, user_id, role (owner/editor/viewer) | Enforces access control |
| Presence | doc_id, user_id, cursor_position, last_heartbeat | Drives real‑time UI cues (who’s online, where they are) |
The ops field in Revision stores a list of low‑level operations (insert, delete, format) that can be replayed to reconstruct any version.
High‑Level API Sketch
POST /documents # create a new doc
GET /documents/{doc_id} # fetch latest snapshot
GET /documents/{doc_id}/rev # fetch specific revision
PATCH /documents/{doc_id} # apply a batch of ops (client‑side edit)
GET /documents/{doc_id}/presence # long‑poll or WebSocket for live cursors
POST /documents/{doc_id}/share # invite users / change permissions
The PATCH endpoint is where concurrency control lives. Clients send a batch of ops together with a base revision number; the server validates and transforms them before persisting.
Architectural Overview (text diagram)
+-------------------+ +-------------------+ +-------------------+
| Client (Web/ |<-----> | Collaboration |<-----> | Persistence |
| Desktop/Mobile) | WS/HTTP| Service (OT/ | RPC | Service (DB) |
+-------------------+ | CRDT Engine) | +-------------------+
^ ^ ^
| | |
| | +--- Indexing Service (search)
| +------- Presence Service (WebSocket)
+----------- Auth & ACL Service
- Collaboration Service: core of real‑time sync, runs the operational transformation (OT) or conflict‑free replicated data type (CRDT) algorithm.
- Persistence Service: stores snapshots and incremental revisions in a durable store (e.g., a document‑oriented DB or a relational DB with JSON columns).
- Presence Service: lightweight in‑memory store (Redis‑like) for broadcasting cursor locations.
- Indexing Service: extracts text for full‑text search, updates asynchronously.
- Auth & ACL Service: validates every request against the Permission table.
Deep Dive: Real‑Time Concurrency Control
Operational Transformation (OT)
- Idea: Transform incoming ops against concurrent ops so that all clients converge to the same state.
- Workflow:
- Client sends ops with a base revision number.
- Server looks up ops that arrived after that base.
- Server transforms the new ops against each later op (using transformation functions
T(insert, delete), etc.). - Server applies the transformed ops, increments the revision, and broadcasts the result.
- Pros: Mature, well‑understood, works well for linear text.
- Cons: Complex to implement correctly; transformation functions must be exhaustive.
Conflict‑Free Replicated Data Types (CRDT)
- Idea: Each client can apply ops locally; merging is deterministic without a central transformer.
- Common choice: Sequence CRDT like RGA or Logoot.
- Pros: Simpler client logic, tolerant to network partitions, no transformation step on the server.
- Cons: Metadata overhead grows with the number of edits; garbage‑collecting tombstones can be tricky.
In an interview, you can argue that OT is a safe default for a product that expects heavy collaboration on large documents, while CRDT shines when you need strong offline support and want to avoid a single point of transformation.
Scaling the System
- Sharding documents: Partition by
doc_idhash; each shard runs its own collaboration instance. This limits the number of concurrent editors per node. - Stateless front‑ends: Load balancers route WebSocket connections to any front‑end; session affinity is optional because state lives in the collaboration service.
- Write‑ahead log: Persist ops to an append‑only log before applying them, enabling replay for recovery.
- Snapshotting: Every N revisions, create a full snapshot to bound replay time for new readers.
- Cache hot documents: Keep recent snapshots in a fast cache (e.g., Redis) to serve read‑only viewers instantly.
Trade‑offs and Design Decisions
| Decision | Option A | Option B | When to pick A |
|---|---|---|---|
| Concurrency algorithm | OT | CRDT | Expect many simultaneous editors, want proven latency bounds |
| Storage model | Document DB (e.g., Mongo) | Relational DB with JSON | Need flexible schema for ops vs. strong transactional guarantees |
| Consistency model | Strong (synchronous) for active editors | Eventual for offline sync | If offline editing is a core use‑case, favor eventual consistency |
| Search indexing | Real‑time incremental indexing | Batch nightly indexing | Real‑time search needed for large enterprises |
Explain that each choice impacts latency, storage cost, and implementation complexity. Interviewers love to see you weigh them against the product’s priorities.
Common Follow‑Up Questions
- How would you handle large binary objects (images) embedded in a document?
- Store them in an object store (S3‑compatible) and reference via URLs. Use a CDN for fast delivery. Keep metadata (size, mime) in the document’s ops list.
- What happens if two users edit the same character at the same time?
- OT resolves by defining a deterministic tie‑breaker (e.g., user ID order). CRDT merges by assigning unique identifiers to each character.
- How do you enforce access control in real‑time streams?
- Every message passes through the Auth & ACL service; the presence service only broadcasts to users with a valid
editororviewerrole.
- Every message passes through the Auth & ACL service; the presence service only broadcasts to users with a valid
- Can you support comment threads attached to a range of text?
- Model comments as separate entities linked to a stable identifier (e.g., a position token from the CRDT). Updates to the document automatically adjust the range via the same transformation logic.
- How would you test the system for correctness under concurrent edits?
- Write property‑based tests that generate random edit sequences, apply them in different orders, and assert convergence. Use integration tests with multiple simulated clients.
How to Practice This
- Sketch the diagram on paper – draw the components, label the data flow, and rehearse explaining each piece in under two minutes.
- Write a short answer script – use a template (situation → challenge → solution → impact) but keep it under 90 seconds; practice aloud, perhaps with Call Assistant to keep you on track.
- Implement a mini‑OT or CRDT demo – a few hundred lines in any language helps you speak confidently about transformation functions and edge cases.
FAQ
- Q: Do I need to implement both OT and CRDT? A: No. Choose one based on the interview’s focus. Explain why you prefer one and acknowledge the other’s trade‑offs.
- Q: How much storage does each revision need? A: It depends on the operation granularity; typically a few dozen bytes per edit. Explain that you can compress batches and prune old revisions via snapshotting.
- Q: What latency target should I aim for? A: Sub‑second round‑trip is common for interactive editing; you can mention 200‑300 ms as a realistic target for a well‑engineered stack.
- Q: Is a CDN necessary for document content? A: It’s useful for static assets like images and for serving large documents to geographically dispersed viewers, but not required for the core text sync.
Frequently asked questions
Do I need to implement both OT and CRDT?
No. Pick one based on the interview focus. Explain why you prefer OT for proven low latency or CRDT for strong offline support, and acknowledge the other’s trade‑offs.
How much storage does each revision need?
It varies with operation size, but a typical edit is a few dozen bytes. Mention batching, compression, and snapshotting to keep storage bounded.
What latency target should I aim for?
Sub‑second round‑trip is the norm for interactive editors; citing 200‑300 ms 99th‑percentile latency shows you understand user expectations.
Is a CDN necessary for document content?
A CDN helps deliver static assets like embedded images and can reduce load on the main service, but the core text sync works without it.
#system design#collaborative editor#real-time sync#concurrency control#scalability#a collaborative document editor like Google Docs