Every product eventually sends notifications: a push when an order ships, an email receipt, an SMS one-time code, a badge in an in-app inbox. The first version is usually a function call inside the feature that needs it. That works until the third team adds one, users start getting the same alert twice, a marketing campaign drowns out a password-reset code, and nobody can answer the question "did this user get told?"
A notification system is the shared service that answers that question. It decides whether to send at all, on which channel, when and in which language, and it remembers what it sent. This article builds it from first principles: the event contract, dedupe and throttling, templates, provider adapters and the feedback loop that keeps addresses clean.
What the system has to guarantee
Write the guarantees down before choosing any technology, because they decide the design. A useful set is:
- No duplicates the user can see. Producers retry, queues redeliver, and workers crash after sending but before acknowledging. Each of these can cause a second send unless every stage is idempotent.
- Consent is enforced centrally. If a user turned off promotional email, no team can bypass that by calling the email provider directly. Legal requirements for marketing consent vary by country, so the rules belong in one place.
- Critical messages are never starved. A one-time login code waiting behind ten million campaign emails is an outage, even though every component is healthy.
- Every message has a traceable state. Support should be able to look up a user and see accepted, suppressed (and why), sent, delivered or failed for each message.
- Delivery is at least once inside the system, effectively once at the edge. True exactly-once delivery to a phone is impossible because the provider call and your bookkeeping cannot be one transaction. You get close by deduplicating on keys at every hop and by using provider-side collapse where it exists.
Guaranteed delivery to the device is not on the list: push services are best-effort, so important information should also live in an in-app inbox the client can fetch.
The pipeline at a glance
The key design choice is that producers emit events ("order o_5521 shipped"), not messages ("send this push with this text"). The orchestrator owns the decision of what to send. That lets you change channels, copy and rules without touching the producing service, and it keeps one place that can enforce consent and dedupe.
Intake validates the event, records it on a durable log and returns. Everything after that is asynchronous. Each channel has its own queues and workers, so an email outage does not slow push, and each lane has a transactional queue (codes, alerts, receipts) separate from a bulk queue (digests, campaigns).
The event contract and the catalog
A small, strict contract avoids most later problems. The event carries a stable ID generated by the producer, a type, the recipient, parameters for the template, and an expiry:
{
"event_id": "01J9ZK4Q2W7M3X", // producer-generated, stable across retries
"type": "order.shipped", // looked up in the notification catalog
"user_id": "u_83121",
"occurred_at": "2026-10-01T07:02:11Z",
"expires_at": "2026-10-02T07:02:11Z", // after this, do not send at all
"params": { "order_id": "o_5521", "carrier": "DHL", "eta": "2026-10-03" },
"collapse": "order:o_5521:shipping" // optional override of the catalog rule
}The expiry matters: a "driver arriving" push delivered forty minutes late is worse than none, so workers drop expired messages and record them as expired.
The catalog is configuration, reviewed like code, that maps each event type to channels, priority, templates, collapse rule, dedupe window and preference category. Producers cannot send a type that is not in it, which gives product and legal teams one list of everything you send.
Preferences, consent and quiet hours
Preferences are a per-user matrix of category by channel. Store explicit rows, default from the catalog when no row exists, and record where each consent came from and when, for example "promotions / email / granted / signup form / 2026-03-02". That audit trail answers "why did I get this?"
Some categories, such as security alerts, cannot be turned off; mark them in the catalog rather than in worker code. Addresses carry state too: verified, bouncing or suppressed. Send SMS only to verified numbers, or a typo sends someone's codes to a stranger.
Quiet hours need the user's time zone, which you should store explicitly and not infer from the IP address on every request. The rule "no non-critical push between 22:00 and 08:00 local time" is applied by computing a send-at time rather than dropping the message. The scheduler holds it until then and the expiry still applies. Critical messages ignore quiet hours, because a fraud alert at 02:00 is exactly when the user needs it.
Dedupe at two levels, collapse at the edge
There are two different duplicate problems, and they need different keys.
Mechanical duplicates come from retries and redelivery. The fix is an idempotency key: the event ID at intake, and event ID plus channel at send time, claimed with an atomic add-if-absent such as Redis SET key 1 NX EX 86400. If the claim fails, another worker already owns the message.
Semantic duplicates are distinct events that mean the same thing, such as two services both emitting "payment failed" for one invoice. Here the key is built from the subject: user, collapse group and channel, with a window from the catalog. The first event in the window wins.
Providers add a third tool. APNs accepts an apns-collapse-id header of up to 64 bytes; a new notification replaces an undelivered one with the same ID. FCM has a collapse key for the same purpose. Use the collapse group as the value, so an offline phone shows only the latest update. In pseudocode:
def handle(event):
# 1. Idempotent intake: a retried producer call must not notify twice.
if not dedupe.add(f"evt:{event.id}", ttl=86400):
return "duplicate_event"
spec = catalog[event.type] # channels, priority, collapse rule, template ids
prefs = preferences.get(event.user_id)
for channel in spec.channels:
if not prefs.allows(event.type, channel): # consent and category opt-outs
log(event, channel, "suppressed_pref"); continue
if channel == "sms" and not prefs.sms_verified:
log(event, channel, "suppressed_unverified"); continue
# 2. Semantic dedupe: same thing about the same object inside the window.
key = f"sem:{event.user_id}:{spec.collapse_key(event)}:{channel}"
if not dedupe.add(key, ttl=spec.dedupe_window_s):
log(event, channel, "suppressed_dupe"); continue
# 3. Per-user cap, except for transactional security messages.
if spec.priority != "critical" and not caps.take(event.user_id, channel):
digest.add(event.user_id, channel, event); continue
# 4. Quiet hours move the send time; they never drop critical messages.
send_at = prefs.next_allowed_time(channel, now(), spec.priority)
queue_for(channel, spec.priority).enqueue(
Message(event.id, event.user_id, channel, spec.template(channel),
event.params, send_at, idempotency_key=f"{event.id}:{channel}"))
Throttling, priority and digests
Provider limits protect your account: exceeding a provider's send rate causes rejections or reputation damage. Each adapter runs a token bucket sized to the provider contract and backs off on throttling responses.
User limits protect attention. A per-user, per-channel cap, for example three non-critical pushes per hour, stops a busy document from becoming fifty pings. Messages over the cap go into a digest bucket, and the scheduler later sends one summary ("12 new comments on 3 documents").
Enforce priority structurally: the transactional queue gets its own workers and provider allowance, which campaigns never borrow. Large fan-outs go through a batch path that expands the audience in chunks and can be paused as a whole.
Templates and localisation
Templates are versioned artefacts, one per event type, channel and locale, with a fallback chain such as pt-BR -> pt -> en. Render at send time, not at intake, so a message delayed by quiet hours uses the current template and the user's current language. Escape every parameter for the target format: HTML-escape for email bodies, and plain text with a length check for SMS and push.
Each channel has hard size limits that the renderer must check before calling the provider. Apple documents a 4096-byte limit for a regular remote notification payload, and FCM enforces its own payload limit, which you should check in Firebase's current documentation. A product name in a long-script language can push a template over the limit, so truncate fields deliberately rather than let the provider reject the message. In CI, render every template in every locale and fail on missing variables and oversized payloads.
Provider adapters and token hygiene
Each channel adapter turns a rendered message into a provider call and classifies the response into one of four outcomes: accepted, retry later (timeouts, 5xx, throttling, with exponential backoff and jitter), permanent failure for this message (a payload error, which goes to a dead-letter queue for a human to inspect), and permanent failure for this address, which feeds the suppression list.
The last outcome is the most important for long-term health. For APNs, an HTTP 410 response means the device token is no longer active for that topic. Delete it and stop sending. For FCM's HTTP v1 API, Firebase's guidance is that UNREGISTERED (HTTP 404) means the registration will never be valid again, and that INVALID_ARGUMENT (HTTP 400) also justifies deletion, but only when you are certain the payload itself is valid. Otherwise a template bug would wipe out good tokens. Firebase also notes that by default it treats a registration as stale if the app instance has not connected for a month, and that on Android a registration inactive for 270 days is expired and garbage-collected.
Store tokens with a last-refreshed timestamp, have the app re-register on start, and prune stale ones. Email has the same loop: hard bounces and complaints must suppress the address.
Worked example: one shipment, three updates, a sleeping user
A user in Lisbon has push enabled, promotional email off and quiet hours from 22:00 to 08:00. At 23:10 the warehouse service emits order.shipped for order o_5521. Its first attempt times out, so it resends the same event ID.
- Intake accepts the first copy and claims
evt:01J9ZK4Q2W7M3X. The retry fails the claim and is recorded as a duplicate event. One event continues. - The catalog says
order.shippedgoes to push and the in-app inbox, priority normal, collapse grouporder:o_5521:shipping, dedupe window 30 minutes. - The inbox entry is written immediately. Quiet hours apply to push, so the scheduler sets send-at to 08:00.
- At 23:25 and 01:15 the carrier emits two
order.in_transitupdates in the same collapse group. The 23:25 update falls inside the 30-minute window and is suppressed on both channels as a semantic duplicate. The 01:15 update is outside it, so it adds an inbox entry, and because the scheduler keeps at most one pending push per collapse group, it replaces the text of the 08:00 push. - At 08:00 the push worker claims
send:{event}:push, renders the Portuguese template, setsapns-collapse-idto the collapse group and sends. APNs accepts it. Delivery status is recorded.
The user wakes to one push with the latest status, and two inbox entries if they want the history. Without the pipeline they would have received four pushes overnight.
Failure modes
- Send then crash. The worker sends, then dies before recording success. The redelivered message hits the send claim and is skipped. The cost is that a crash between claim and send loses that message, so expire stale claims and reconcile against provider responses where you can.
- Backlog replay. After an outage, stale messages drain at once. Expiry and per-user caps stop the flood.
- Priority inversion. A campaign shares a provider rate limit with one-time codes and the codes time out. Use separate allowances, and if possible separate sender identities.
- Template regression. A deploy renders a missing variable to every user. Canary new template versions on a small share of traffic and watch rendering errors.
- Silent token decay. The device-token table keeps growing, and the delivered-to-sent ratio falls month by month. Watch that ratio per platform and prune.
Operational guidance and trade-offs
Track per event type and channel: accepted, suppressed by reason, sent, failed by class, and latency from event to provider acceptance. Alert on transactional latency, not averages. Give support a searchable per-user message log.
Trade-offs: an orchestrator adds a hop, but it is the only way to enforce consent and dedupe across products. Longer dedupe windows remove noise but can hide a real second failed payment. Caps protect attention but delay information. Push is fast but unreliable, email reliable but slow to be read, SMS reliable and expensive, so choose channels per event type.
For the reliable handoff from the producing service's database to the intake log, use the transactional outbox pattern. For idempotency keys in general see idempotency, and for poison messages see dead letter queues. If you also deliver events to partner systems, webhook delivery shares much of this retry and signing machinery. For email reputation in depth see email delivery and spam filtering.
What to do next
- List every notification your product sends today and turn the list into a catalog with type, channels, priority, collapse rule and preference category.
- Change one producer to emit an event with a stable ID and an expiry instead of calling a provider directly.
- Add idempotent intake and a send-time claim keyed on event ID and channel, and test them by replaying a batch of events twice.
- Store preferences with consent source and timestamp, store each user's time zone, and implement quiet hours as a send-at time.
- Split transactional and bulk traffic into separate queues with separate worker pools and provider allowances.
- Handle APNs 410 and FCM UNREGISTERED in your adapters, add a last-refreshed timestamp to device tokens, and put the delivered-to-sent ratio on a dashboard.
- Add a template test that renders every locale and fails on missing variables and payloads over the provider limits.