AI System Design
← All chapters
Chapter 10 6 min read

Design a Notification System

A notification system delivers push, SMS, and email messages on behalf of many internal services through third-party providers like APNs, FCM, Twilio, and SendGrid. The design moves from a fragile single server to a decoupled pipeline of per-channel queues and workers, then adds the reliability, deduplication, rate limiting, and monitoring that make it trustworthy.

Architecture at a glance
  1. Service (caller)
  2. Notification servers (auth, validate)
  3. Cache / DB (user, device, templates)
  4. Per-channel message queues
  5. Workers
  6. Third-party provider (APNs / FCM / SMS / email)
  7. User device

Problem & requirements

Many parts of a product need to tell users something: a payment cleared, a friend replied, a flight changed. Rather than each team integrating with Apple, Google, and SMS vendors separately, a shared notification platform accepts requests and handles delivery. Clarify which channels are needed, whether delivery must be real time or can tolerate small delays, which devices are supported, and whether users can opt out.

A reasonable scope is iOS push, Android push, SMS, and email; soft real time, meaning seconds of delay under load is acceptable; and support for user opt-out. Notifications can be triggered by client apps or scheduled on the server.

Back-of-the-envelope estimation

A typical volume is 10 million mobile pushes, 1 million SMS, and 5 million emails per day, about 16 million in total. Averaged over 86,400 seconds that is about 185 notifications per second; pushes alone are 10M / 86,400 ≈ 116 per second. Traffic is very bursty, though: a marketing campaign or breaking event can send millions within minutes, so design for peaks one or two orders of magnitude above the average.

Payloads are small, perhaps 1–5 KB including template data, so 16M × 5 KB = 80 GB per day of notification log at most. The bottleneck is not storage but third-party throughput limits and the cost of SMS, which is why rate limiting and channel choice matter.

Notification types and contact information

Each channel has its own path. iOS push needs a provider (your server) to build a request with the device token and a JSON payload and send it to Apple Push Notification service (APNs), which delivers to the device. Android uses Firebase Cloud Messaging (FCM) in the same role. SMS goes through vendors such as Twilio or Nexmo, and email through providers such as SendGrid or Mailchimp, because running your own mail servers with good deliverability is hard.

To send anything you need contact details. When a user installs the app or signs up, the API servers record their email and phone in a user table and their device tokens in a device table. One user can have many devices, so the device table is one-to-many, and stale tokens must be removed when providers report them invalid.

From a single server to a decoupled design

A first draft puts one notification server between callers and providers. It works, but it has three problems: it is a single point of failure; it is hard to scale because the database, cache, and processing all live together; and it is a performance bottleneck because building HTML emails and waiting on slow third-party responses block everything else.

The improved design separates concerns. Notification servers become a stateless, horizontally scaled tier that authenticates callers, validates requests, fetches user and template data from the cache or database, and drops a message onto a queue. There is one message queue per channel, so an outage at the SMS vendor fills only the SMS queue while pushes and emails keep flowing. Workers pull from each queue and call the matching provider. Each tier can now scale on its own.

  • Notification servers: auth, validation, rate limits, enqueue.
  • Queues: buffer bursts and isolate channel failures.
  • Workers: talk to providers, retry, record results.
Figure 1One queue per channel
One queue per channelTHIRD-PARTY PROVIDERSsendlookupCaller servicesbilling, shipping, cronNotification serversauth, validate, enqueueUser + device cachetokens, phones, emailsPush queueSMS queueEmail queuePush workersSMS workersEmail workersAPNsiOSFCMAndroidSMS providerEmail providerSMTP / API
Notification servers only validate and enqueue, so slow providers never block callers. Because each channel has its own queue and workers, an SMS vendor outage backs up only the SMS lane while pushes and emails keep flowing.

Reliability: no data loss and deduplication

Notifications may be delayed or reordered, but they should not vanish. The system persists each notification in a notification log database before or as it enqueues it, and workers update its status. If a worker crashes mid-send, the record shows the notification as unsent, and a retry job can pick it up.

Exactly-once delivery is impossible across a network you do not control, because a provider may accept a message and then the acknowledgement may be lost. So the system aims for at-least-once delivery plus deduplication. Each event carries an ID; before sending, a worker checks whether that ID has already been processed and skips it if so. This gives behaviour that is close to exactly once in practice, while accepting rare duplicates in edge cases.

Figure 2Logged, retried, deduplicated
Logged, retried, deduplicatedNotif serverNotif logPush queueWorkerAPNs1. insert event e91, status = pending2. enqueue e91 (token, payload)3. deliver e914. already sent e91?5. no, still pending6. send push7. timeout8. re-enqueue e91 with backoff9. redeliver e91 after delay10. send push11. 200 accepted12. mark e91 sent
The notification log is written before enqueueing, so a crashed worker leaves a pending record a retry can find. The dedup check on event_id turns at-least-once redelivery into near exactly-once sends.

Templates, settings, rate limiting and retries

Templates keep formatting consistent and save callers from building full messages: a template has placeholders for name, item, or date, and the server fills them in. Notification settings store, per user and channel, whether they opted in; workers check these before sending, both for user experience and for legal compliance on marketing messages.

Rate limiting caps how many notifications a user receives in a period, because users who feel spammed turn notifications off entirely, which is far worse than one missed message. Retries with backoff handle transient provider failures; if a provider keeps failing, alert operators and consider failing over to a secondary vendor.

Figure 3Gates before the queue
Gates before the queuereadopted inopted outover limitallowedenqueueSend requestuser, template_id, varsOpt-in checkper user, per channelRate limitermax N per user per dayRender templatefill vars, pick localeSettings DBuser × channel → on/offDroppedlogged, never sentChannel queue
Cheap checks run first: a user who opted out of a channel, or has hit the per-user cap, is dropped before any template is rendered or any provider quota is spent.

Security, monitoring and tracking

Only authorised internal clients should be able to send notifications, otherwise the platform becomes a spam cannon. Callers authenticate with an appKey and appSecret (or an equivalent token scheme), and the notification servers reject anything unsigned.

The most important operational metric is queue depth: a growing backlog means workers cannot keep up and more should be added. Beyond that, event tracking records open rates, click rates, and engagement by integrating with the analytics service, which lets product teams judge whether a notification was worth sending. A notification moves through states such as start, pending, sent, delivered, clicked, or error, and tracking those transitions makes delivery problems visible.

Trade-offs & wrap-up

The final design trades simplicity for isolation and durability. Queues add latency and operational overhead, but without them one slow provider would stall all traffic. Persisting a log before sending costs a write per message but buys recoverability. Deduplication costs a lookup per message but keeps users from seeing the same alert twice.

In an interview, walk through the evolution explicitly: name each weakness of the single-server design and the component that removes it. Then cover reliability, settings, rate limits, retries, security, and monitoring as the features that turn a working prototype into a platform other teams can depend on.

Key numbers

Push per day
10 million (≈116/s avg)
SMS per day
1 million
Email per day
5 million
Total average rate
≈ 185/s, with large bursts
Max log volume
≈ 80 GB/day at 5 KB each

Key terms

APNs
Apple Push Notification service, the gateway that delivers push messages to iOS devices.
FCM
Firebase Cloud Messaging, Google's service for delivering push messages to Android (and other) devices.
Device token
A unique identifier issued by the push provider that addresses one app install on one device.
Notification log
A durable record of every notification and its status, used for retries and auditing.
At-least-once delivery
A guarantee that a message is delivered one or more times, never zero, with duplicates possible.
Deduplication
Skipping events whose ID has already been processed so retries do not produce repeat notifications.
Queue depth
The number of messages waiting in a queue, the main signal that workers are falling behind.

Common mistakes

  • Using one shared queue for all channels so an SMS vendor outage blocks push and email.
  • Promising exactly-once delivery instead of at-least-once with deduplication.
  • Calling third-party providers synchronously from the request path.
  • Forgetting user opt-out settings and per-user rate limits.
  • Leaving the send API unauthenticated, letting any service spam users.
  • Not monitoring queue depth, so backlogs are discovered by users rather than alerts.

Further study

  • Apple Push Notification service documentation
  • Firebase Cloud Messaging architectural overview
  • Netflix TechBlog: Rapid Event Notification System (RENO)
  • LinkedIn Air Traffic Controller (notification volume control)

Now practise it