When an interviewer asks you to design a notification service, they’re looking for how you translate a simple‑looking problem into a robust, production‑grade system. The goal is to send messages (email, SMS, push, in‑app) to millions of users, possibly with different priorities, while keeping latency low and reliability high. Below is a practical walkthrough you can use in a real interview.

1. Clarify the Scope First

Start by asking clarifying questions. The answers will shape the rest of your design.

  • Types of notifications – Are we supporting only push and email, or also SMS, webhooks, etc.?\
  • Target audience – Is it a single product’s users, or a platform serving many downstream services?\
  • Volume expectations – Roughly how many notifications per second during a peak? You can respond with a range (e.g., "tens of thousands to a few hundred thousand per second") without naming exact numbers.

Typical functional requirements include:

  • Send a notification (create request).
  • Track delivery status (sent, delivered, failed).
  • Support retries with back‑off.
  • Rate‑limit per user or per channel to avoid spamming.
  • Provide a simple query API for status and history.

Non‑functional requirements often surface:

  • Low latency – end‑to‑end < 1 s for most channels.
  • High availability – tolerate node or datacenter failures.
  • Scalability – handle traffic spikes without degradation.
  • Durability – no loss of notification data.
  • Observability – metrics, logs, and alerts for failures.

2. Identify Core Entities and Data Model

A clean data model helps you reason about storage and APIs.

EntityKey FieldsPurpose
Notificationid, user_id, channel, payload, priority, created_at, statusRepresents a single message to be delivered.
UserChannelConfiguser_id, channel, enabled, preferencesStores per‑user opt‑in/opt‑out and throttling settings.
DeliveryAttemptid, notification_id, attempt_num, timestamp, result, error_codeTracks each try to deliver the message.
RateLimitBucketuser_id, channel, tokens, last_refillImplements token‑bucket rate limiting.

You don’t need a fully normalized schema; a denormalized document store (e.g., DynamoDB, MongoDB) works well for fast writes and reads.

3. Define the Public API

Keep the API surface minimal but expressive.

POST /notifications
{
  "user_id": "12345",
  "channel": "email",
  "payload": {"subject": "Welcome", "body": "Hello!"},
  "priority": "high"
}
GET /notifications/{id}
Response: {"status": "delivered", "attempts": 1}
GET /users/{user_id}/notifications?limit=50&offset=0
Response: [{"id": "...", "channel": "push", "status": "failed"}, ...]

These endpoints let the caller create a notification, query its status, and list recent messages.

4. High‑Level Architecture

A text diagram works well on a whiteboard:

Client → API Gateway → Ingestion Service → Message Queue → Worker Pool → Channel Adapters → Delivery Providers
        │                 │                │                │                │
        └─ Auth Layer ────┘                └─ Retry Store ──┘

4.1 Ingestion Service

  • Validates request, checks user preferences, and writes a Notification record.
  • Enqueues a message (e.g., Kafka topic) containing the notification id.
  • Idempotency – generate a client‑provided idempotency key to avoid duplicate entries.

4.2 Worker Pool

  • Consumers read from the queue, fetch the full notification payload, and apply rate limiting.
  • Retry logic lives here; failed attempts are re‑queued with exponential back‑off.
  • Parallelism – scale workers horizontally; each worker is stateless.

4.3 Channel Adapters

Separate adapters for each channel (email, push, SMS). They abstract away provider‑specific APIs (SendGrid, APNs, Twilio, etc.).

  • Circuit breaker per provider to stop hammering a flaky endpoint.
  • Metrics per channel (success rate, latency).

4.4 Persistence Layer

  • Primary store – a durable NoSQL table for Notification and DeliveryAttempt.
  • Secondary indexes for fast lookup by user and status.
  • Write‑ahead log (WAL) or transaction log for recovery.

5. Deep Dive: Hard Parts

5.1 Ordering Guarantees

If a user receives multiple notifications, they often expect order preservation per channel. Use a partition key (user_id + channel) in the queue so all messages for that user/channel land in the same partition, preserving order. Trade‑off: this can create hot partitions for very active users; you may shard further by time bucket.

5.2 Rate Limiting & Throttling

Implement a token‑bucket per user/channel. Workers check the bucket before sending; if empty, they defer the message. Store bucket state in a fast in‑memory store (Redis) with TTL for automatic refill.

5.3 Failure Handling & Retries

  • Transient failures (network hiccups) – retry with exponential back‑off up to a max attempts count.
  • Permanent failures (invalid address) – mark notification as failed and stop retrying.
  • Dead‑letter queue – capture messages that exceed retry limits for later analysis.

5.4 Multi‑Region Resilience

Deploy the ingestion service and queue in multiple regions. Use a global load balancer to route clients to the nearest region. Replicate the primary store asynchronously; if a region goes down, traffic fails over to another region and the queue buffers the backlog.

6. Trade‑offs and Alternatives

DecisionProCon
Single queue vs. per‑channel queuesSimpler architecture, easier to reason about ordering.May cause a single point of contention; per‑channel queues allow independent scaling.
NoSQL store vs. relational DBFast writes, easy horizontal scaling.Harder to enforce complex constraints; eventual consistency for some queries.
Push‑only vs. Pull‑based deliveryPush reduces latency for real‑time channels.Pull (e.g., email batch) can be more cost‑effective for low‑priority traffic.

Explain why you pick one approach, then acknowledge the alternative and when you’d switch.

7. Follow‑Up Questions Interviewers May Ask

  1. How would you handle a burst of 10× traffic for a single user? – Talk about sharding the partition key, burst‑capacity buffers, and back‑pressure to the ingestion layer.
  2. What if a provider’s API is rate‑limited globally? – Show circuit‑breaker logic, fallback providers, and graceful degradation (e.g., delay low‑priority messages).
  3. How do you guarantee exactly‑once delivery? – Discuss idempotent writes, deduplication keys, and the trade‑off between exactly‑once and at‑least‑once semantics.
  4. How would you monitor SLA compliance? – Mention metrics (latency, success rate), dashboards, alert thresholds, and automated health checks.
  5. Can you add support for user‑specific scheduling (send at 9 am)? – Add a scheduled_at field, store in a time‑ordered table, and have a scheduler worker pull due notifications.

8. Where Call Assistant Helps You Practice

When you rehearse your answer, a tool like Call Assistant can listen to your mock interview, surface the next logical question, and let you keep the conversation on track. It also records a concise version of your story grounded in your resume, so you can iterate quickly.

9. How to Practice This

  1. Mock the interview – Pair with a peer, use a timer, and run through the entire flow, including follow‑up questions.
  2. Sketch the diagram on paper – Focus on labeling components and data flow; avoid perfection, aim for clarity.
  3. Write a short script – Record yourself answering the API and hard‑part sections, then listen back (or use Call Assistant) to trim filler and tighten language.

FAQ

  • What is the simplest way to store notification data? Use a document‑oriented NoSQL table where each item contains the full payload and status fields. It gives fast writes and easy scaling without complex joins.

  • How do you ensure low latency for push notifications? Keep the push path short: ingest → queue → worker → push adapter → provider. Use in‑memory caches for rate‑limit checks and batch small payloads to reduce round‑trips.

  • When should you prefer a pull‑based delivery model? For bulk or low‑priority channels like email newsletters, pulling batches from a store lets you control send rates and reduce provider costs.

  • What metrics matter most for a notification service? Delivery latency, success rate per channel, retry count, and queue depth. Monitoring these helps you spot bottlenecks before they affect SLAs.

Frequently asked questions

What is the simplest way to store notification data?

Use a document‑oriented NoSQL table where each item contains the full payload and status fields. It gives fast writes and easy scaling without complex joins.

How do you ensure low latency for push notifications?

Keep the push path short: ingest → queue → worker → push adapter → provider. Use in‑memory caches for rate‑limit checks and batch small payloads to reduce round‑trips.

When should you prefer a pull‑based delivery model?

For bulk or low‑priority channels like email newsletters, pulling batches from a store lets you control send rates and reduce provider costs.

What metrics matter most for a notification service?

Delivery latency, success rate per channel, retry count, and queue depth. Monitoring these helps you spot bottlenecks before they affect SLAs.

#system design#notification service#architecture#scalability#interview prep#a notification service