When an interviewer asks you to design a ticket booking system, they’re looking for how you translate a real‑world product into a clean, scalable architecture. The problem is deceptively simple – users pick an event, select seats, pay, and receive a confirmation. The difficulty lies in handling high traffic spikes, preventing double‑booking, and keeping latency low for a smooth user experience.
1. Clarify the Scope
Start by asking clarifying questions. Typical points include:
- Event types – concerts, movies, sports? Different events may have different seat layouts and pricing rules.
- Ticket volume – are we designing for a regional theater or a global platform that sells millions of tickets per day?
- User flow – do we need to support seat‑selection, or is it a simple "any available" model?
- Payment – will the system integrate with external payment gateways, and do we need to store payment details?
- Availability window – does the system need to hold seats for a short period while the user checks out?
- SLAs – what latency targets are expected for the search and checkout steps?
These questions guide the functional and non‑functional requirements you’ll capture next.
2. Functional & Non‑Functional Requirements
Functional
- Event catalog – create, update, and retire events.
- Seat inventory – expose real‑time seat availability per event.
- Reservation workflow – hold seats, confirm purchase, release on timeout.
- Payment processing – invoke external gateway, handle callbacks.
- Order history – let users view past purchases and download tickets.
- Cancellation & refunds – support policy‑driven refunds.
Non‑Functional
- Scalability – handle traffic spikes (e.g., ticket releases for popular concerts).
- Low latency – sub‑second response for seat lookup and checkout.
- Strong consistency – avoid double‑booking, especially during high contention.
- Availability – graceful degradation; the system should stay functional even if a downstream service fails.
- Security & compliance – protect payment data and personal information.
- Observability – metrics, tracing, and alerts for latency and error rates.
3. Core Entities & Data Model
| Entity | Key Attributes | Relationships |
|---|---|---|
| Event | eventId, name, venueId, startTime, endTime, pricingRules | belongs to Venue; has many Seats |
| Venue | venueId, name, address, seatMap | contains Seats |
| Seat | seatId, venueId, section, row, number, priceTier | belongs to Venue |
| Inventory | eventId, seatId, status (available, held, sold) | maps Seat to Event |
| Order | orderId, userId, eventId, seatIds, totalAmount, status, createdAt | references Event, Seats |
| Payment | paymentId, orderId, gateway, status, transactionId | belongs to Order |
The Inventory table is the heart of the system – it tracks the current state of each seat for a given event.
4. API Surface
4.1 Public REST‑like endpoints
GET /events?city=NYC&date=2026-10-01 # List events
GET /events/{eventId}/seats?section=A # Seat availability
POST /orders # Create reservation (holds seats)
POST /orders/{orderId}/confirm # Pay and finalize
GET /orders/{orderId} # Order status
DELETE /orders/{orderId} # Cancel (release held seats)
4.2 Internal service contracts
- SeatService – expose
checkAvailability(eventId, seatIds)andholdSeats(eventId, seatIds, ttl). - PaymentService –
initiate(paymentInfo)andhandleCallback(callbackPayload). - NotificationService – send email/SMS confirmations.
5. High‑Level Architecture
+----------------------+ +----------------------+ +-------------------+
| Client (Web/Mobile) |<---->| API Gateway (TLS) |<---->| Load Balancer |
+----------------------+ +----------------------+ +-------------------+
| |
v v
+--------------+ +--------------+
| Front‑End | | Auth/Authz |
+--------------+ +--------------+
| |
v v
+-------------------------------+
| Service Layer (Stateless) |
| - EventService |
| - SeatService |
| - OrderService |
| - PaymentAdapter |
+-------------------------------+
| |
+--------------------------+ +--------------------------+
| |
v v
+----------------------+ +----------------------+
| Cache (Redis) | | Relational DB (Postgres) |
+----------------------+ +----------------------+
| |
v v
+----------------------+ +----------------------+
| Message Queue (Kafka) | | Search Engine (Elastic) |
+----------------------+ +----------------------+
- API Gateway terminates TLS, does request routing, and enforces rate limits.
- Service Layer is stateless; scaling is achieved by adding more instances behind the load balancer.
- Cache stores hot seat‑availability queries and recent event metadata. Cache invalidation occurs on every hold or purchase.
- Relational DB holds the canonical source of truth for inventory and orders. Use row‑level locking or optimistic concurrency to guarantee consistency.
- Message Queue decouples order creation from downstream actions like sending emails or updating analytics.
- Search Engine powers flexible event search (by city, date, genre) without hitting the primary DB.
6. Deep Dive: Concurrency Control (Hard Part)
6.1 The Double‑Booking Problem
When many users request the same seat at the same moment, naïve reads can lead to two orders thinking the seat is free. The system must enforce mutual exclusion.
6.2 Approaches
| Technique | Pros | Cons |
|---|---|---|
| Pessimistic lock (row‑level) | Guarantees exclusive access; simple to reason about | Can become a bottleneck under high contention; holds DB resources longer |
| Optimistic lock (version column) | Low contention cost; fits well with stateless services | Requires retry logic; may waste work if conflicts are frequent |
| Distributed lock (Redis SETNX) | Fast in‑memory lock; works across service instances | Needs careful TTL handling; risk of split‑brain if lock expires prematurely |
A common hybrid solution is to attempt a fast Redis lock first. If the lock is acquired, the service writes a provisional hold record to the DB with a short TTL (e.g., 5 minutes). If the DB write fails because the seat is already held, the service releases the Redis lock and returns a conflict to the client. This pattern keeps the critical section short and reduces DB lock time.
6.3 Holding Seats
- Hold record:
orderId, seatId, heldAt, expiresAt. - A background worker periodically scans for expired holds and releases them back to the pool.
- The hold TTL is communicated to the client so the UI can show a countdown.
7. Deep Dive: Caching Strategy
Seat availability is read‑heavy. Cache the result of GET /events/{eventId}/seats with a short TTL (e.g., 2 seconds) and invalidate the entry on every successful hold or purchase. Use Cache‑Aside: the service first checks Redis; on a miss, it reads from the DB, populates the cache, and returns the data.
For search, a search index (Elastic) holds event metadata. Updates to the event catalog are streamed via the message queue to keep the index eventually consistent.
8. Trade‑offs & Extensions
| Decision | Reasoning |
|---|---|
| Stateless services | Easier horizontal scaling; fits cloud‑native deployment. |
| Relational DB for inventory | Strong consistency guarantees needed for seat state. |
| Redis for short‑lived locks | Low latency, but must guard against lock expiration causing race conditions. |
| Message queue for eventual tasks | Decouples order confirmation from email sending, improving responsiveness. |
Possible Extensions
- Dynamic pricing – adjust seat price based on demand; would require a pricing service and additional data pipelines.
- Seat recommendation – use ML to suggest seats; adds a recommendation microservice.
- Multi‑currency support – introduces currency conversion and compliance layers.
- Mobile push notifications – integrate with a notification hub.
9. Follow‑up Questions Interviewers May Ask
- How would you handle a sudden traffic spike when tickets go on sale?
- Discuss auto‑scaling groups, warm‑up of caches, and pre‑warming the database connection pool.
- What if the payment gateway is slow or fails?
- Explain a saga pattern: mark the order as pending, retry the payment asynchronously, and cancel the hold if the saga aborts.
- How do you ensure data durability for orders?
- Use write‑ahead logging, commit the order and payment records in a single transaction, and replicate the DB across availability zones.
- Can you make the system eventually consistent for seat availability?
- Yes, if we relax the guarantee; we could rely on a distributed lock service and accept brief over‑booking that is later reconciled.
- Where does Call Assistant fit in your interview prep?
- You can practice delivering this design aloud, letting the assistant capture your flow and suggest concise follow‑up answers grounded in your own experience.
10. Sample Answer (45‑90 seconds)
"Sure, I’d start by clarifying the scope: are we building a global platform or a single‑venue system? Assuming a global service, the core problem is keeping seat inventory consistent under high contention. I’d model the main entities—Event, Venue, Seat, Inventory, Order, and Payment. The API would expose endpoints for listing events, checking seat availability, creating a reservation, and confirming payment. Architecturally, I’d use a stateless service layer behind an API gateway, a relational database for the inventory because we need strong consistency, and Redis for fast locks and caching. When a user selects seats, the service attempts a short‑lived Redis lock, writes a provisional hold to the DB with a TTL, and returns a hold token. If the payment succeeds, the hold turns into a sale; otherwise a background worker releases the seats after the TTL. Caching the seat map reduces read load, and a message queue decouples email notifications. Trade‑offs include the extra complexity of managing distributed locks versus the latency benefit they provide. For traffic spikes, I’d rely on auto‑scaling and pre‑warming caches. Finally, I’d instrument latency metrics and set up alerts to keep the checkout experience under a second."
How to practice this
- Sketch the diagram on paper – spend five minutes drawing the components and their interactions without looking at any reference.
- Run a mock interview – ask a friend to play the interviewer, or use Call Assistant to record your answer and get real‑time feedback on pacing and completeness.
- Implement a tiny prototype – code a simple
SeatServicewith an in‑memory lock map; then add unit tests that simulate concurrent reservations to see the race conditions first‑hand.
FAQ
- Q: Do I need a separate microservice for pricing?
- A: Not for the core design. You can embed pricing rules in the Event entity and evolve to a dedicated service later if dynamic pricing becomes a requirement.
- Q: How critical is eventual consistency for seat availability?
- A: For a ticketing system, strong consistency is usually non‑negotiable because double‑booking erodes trust. Eventual consistency can be acceptable for search indexes, not for the inventory itself.
- Q: What are common pitfalls when using Redis locks?
- A: Forgetting to set an appropriate TTL, not handling lock renewal, and relying on a single Redis node without failover can cause deadlocks or split‑brain scenarios.
- Q: Should I store the actual ticket PDF in the database?
- A: It’s better to store a reference (URL) to an object store (e.g., S3) and keep the metadata in the DB. This keeps the DB size manageable and leverages CDN caching for downloads.
Frequently asked questions
What functional requirements are essential for a ticket booking system?
You need an event catalog, real‑time seat inventory, a reservation workflow that holds seats temporarily, payment integration, order history, and cancellation/refund handling.
How do you prevent double‑booking under high concurrency?
Combine a fast Redis lock with a short‑TTL hold record in the relational database. If the DB write fails because the seat is already held, release the Redis lock and return a conflict.
When should you use a cache for seat availability?
Cache the seat map for hot reads, using a short TTL (a few seconds) and invalidate the entry on every successful hold or purchase to keep data fresh.
What trade‑offs exist between pessimistic and optimistic locking?
Pessimistic locking guarantees exclusivity but can become a bottleneck; optimistic locking reduces lock time but requires retry logic when conflicts are frequent.
#system design#ticket booking#concurrency#caching#interview prep#a ticket booking system