Skip to content
Navigation
Dashboard
← All system design problems
Hard · Notification Systems · About 55 minutes

Design a Notification Service - Push, Email, SMS

Understand the requirements. Trace the requests. Explain the trade-offs.

CANDIDATE-LED INTERVIEW WALKTHROUGH

Design a notification service: durable intent, consent and delivery uncertainty

Distinguish a business notification from channel attempts and provider receipts. Explain personalization, suppression, retries, scheduling, quotas and what happens when a provider accepted a message but the response disappeared.

Interviewer replies, workloads and targets are illustrative assumptions to agree on in an interview. Use this as a study resource: establish scope, draw a complete baseline, then choose the most consequential deep dives with your interviewer.

1. Ask what sent means to the product

Candidate explains

When you ask for notifications, do you mean accepting a send request, handing it to a provider, reaching a device, or a user actually reading it? I would establish that vocabulary first. Our system can own durable intent and attempts, but external delivery is only observable to the extent the channel reports it.

Candidate asksIllustrative interviewer replyDesign consequence
Which channels?Push, email, SMS and in-app.One logical notification expands into independently tracked channel deliveries.
Which message types?Transactional alerts and optional campaigns.Separate priority, consent, expiry and throughput policies.
What latency?Transactional provider handoff p99 under one second normally.This excludes carrier delivery and user reading; campaign queues cannot starve transactional traffic.
What volume?100 million logical notifications/day, average 1.2 channel deliveries each.Capacity uses 120M delivery tasks, not 100M provider calls by assumption.
Do users control delivery?Channel/category preferences, quiet hours and unsubscribe.Check current suppression at final dispatch as well as initial expansion.
Are scheduled and bulk sends required?Yes; large campaigns may take minutes.Paginated audience expansion, durable cursors and scheduled eligibility rather than one giant transaction.
Are duplicates acceptable?Avoid them; some providers do not support idempotent sends.Record uncertain outcomes and disclose residual duplicate risk rather than claiming exactly-once delivery.

2. Agree lifecycle and privacy boundaries

Candidate explains

I accept a request only after durable intent. A successful API response means accepted, not delivered. I keep sensitive details out of lock-screen payloads and require authorization when the app fetches the underlying content. Policy rules for message categories need product and compliance review; a priority flag is not permission to bypass consent.

RequirementAgreed target or scopeDesign implication
Acceptance202 with notificationId and status URL after persistence.Caller retries with an idempotency key; same intent is not expanded twice.
PersonalizationVersioned templates, locale fallback and bounded typed variables.Reproducible rendering; escaped content and no arbitrary template code.
PreferencesHonor current suppression and category/channel rules before send authorization.Versioned preference state; a queued task does not permanently freeze consent.
Delivery evidenceTrack accepted, provider accepted, delivered where confirmed, failed, suppressed, expired and unknown.A missing receipt never proves delivery; email delivery does not prove a person read it.
AvailabilityDurable multi-zone intent and isolated channel workers.Provider outage creates bounded backlog, not a blocking request cascade.
RetentionShort-lived content and a separate minimal audit retention policy.Avoid retaining full message bodies and device tokens in every attempt log.

3. Size logical notifications and channel attempts separately

Candidate explains

I use a 5x peak on a 120M-delivery daily workload. Retries, multiple devices and audience expansion can amplify that, so I would measure those distributions rather than multiply only the user count.

QuantityCalculationInterpretation
Logical intent100M/day / 86,400 = 1,157/s average.At 1.2 channels each, there are 120M delivery tasks/day.
Channel tasks120M / 86,400 = 1,389/s average; 6,944/s at 5x.Assume 70% push, 25% email, 5% SMS for this ledger only.
Provider mix84M push, 30M email, 6M SMS/day.Device-level push fan-out can increase the actual attempts beyond these assumed task counts.
Delivery history120M × 500 bytes × 30 days = 1.8 TB logical.Three copies of delivery history = 5.4 TB before indexes. Separate 500-byte attempt rows at 1.1 attempts/task add 1.98 TB logical over 30 days.
Queue envelope120M × 300 bytes = 36 GB/day.Use template and payload references where safe; retries must not duplicate large content blobs.
Provider outageEmail: 30M/day / 24 = 1.25M tasks in a one-hour outage.At 1,000/s recovery capacity minus 347/s new arrivals, drain takes about 32 minutes, excluding retries and quotas.

4. Separate intent, recipient delivery and attempt evidence

Candidate explains

The deduplication boundary is business intent plus recipient plus channel, not a randomly regenerated worker request ID. If a campaign intentionally sends another message tomorrow, that is a new intent. Attempts reuse the delivery identity even when a worker or provider adapter restarts.

Record and keyFields / invariantAccess pattern
Notification(tenantId, notificationId)businessKey, category, templateVersion, payloadRef, audienceRef, expiresAt and priority.Unique caller idempotency key plus payload hash; durable acceptance.
Delivery(notificationId, recipientId, channel)state, nextAttemptAt, preferenceVersion, stable providerKey and lease fence.Unique expansion insert; channel-specific retries and status.
Attempt(deliveryId, attemptNo)provider, requestId, startedAt, result class and receipt reference.Record before/after send; unknown outcome is a real state.
Preference(userId, category, channel)allowed, version, quiet-hour timezone and verified destinations.Authoritative suppression read at dispatch; deletion and invalidation events.
ProviderReceipt(provider, eventId)delivery mapping, timestamp and normalized event type.Deduplicate signed callbacks and tolerate out-of-order events.
InAppInbox(userId, createdAt, deliveryId)authorized content reference, readAt and expiry.Stable cursor pagination; idempotent read-state updates.

5. Design acceptance and provider callbacks

Candidate explains

The external API never lets a tenant send arbitrary templates to another tenant’s users. A provider callback is an untrusted input until its signature, timestamp and account binding are verified. I store receipt evidence before deriving status.

OperationContractFailure or retry behavior
POST /notificationsIdempotency-Key; category, templateVersion, recipient/audience reference, variables, optional schedule. Return 202.Same key with changed payload is 409; invalid template or forbidden audience is rejected before acceptance.
GET /notifications/{id}Tenant-scoped intent status plus aggregate delivery counts.Report partial and unknown results; do not expose recipient PII through a public status URL.
PUT /me/preferencesVersion-aware category/channel updates.Persist suppression before acknowledging; queued dispatches recheck it.
POST /provider-events/{provider}Signed provider event with provider message ID.Deduplicate callback IDs; accept delayed receipts without blindly overwriting newer contradictory evidence.
GET /me/inbox?cursor=...Authenticated in-app inbox with stable ordering.Deleted/expired resources are omitted; a push payload is not authorization to fetch one.

6. Draw intent and each channel boundary

Candidate explains

I separate intake, audience expansion, channel queues, provider adapters and receipt processing. That lets SMS throttling or a push outage leave email and in-app delivery operational. The authoritative ledger connects these asynchronous stages.

7. Trace a request through expansion and consent

Candidate explains

Imagine an order shipment notification. The order service provides a stable shipment-event ID and a template version. We persist that intent once, then create a delivery for each allowed recipient/channel combination. A retry of the order event must not create a second shipment notification.

  1. Accept durable intent

    Validate tenant, category, template and variables; atomically store intent, caller key result and expansion event. Return 202 only after that commit.

  2. Expand in bounded pages

    Persist the audience snapshot or explicitly define dynamic-audience semantics. Workers checkpoint a cursor and insert unique delivery keys so replaying a page is safe.

  3. Apply eligibility

    Check category preferences, verified destination, quiet hours and expiry. Defer to the next eligible time without extending the notification’s business expiry. Suppressed work is terminal unless a new user action creates new intent.

  4. Render reproducibly

    Use immutable template version and locale. Escape variables by output context and avoid storing sensitive content in logs. Expensive rendering can be cached for identical safe inputs.

  5. Queue by priority and channel

    Record channel-ready intent transactionally before publication. Weighted scheduling reserves transactional capacity while preventing indefinite campaign starvation.

8. Explain attempts, receipts and unsubscribe races

Candidate explains

Before calling the provider, the worker validates the current lease, expiry and suppression. To make unsubscribe acknowledgment meaningful, I define the ordering point: preference update versus dispatch authorization is serialized in a per-user authority or transaction. A send already authorized may still be in flight; we cannot recall a delivered SMS.

  1. Respect provider quotas

    Acquire channel/provider capacity before the send, not after accumulating unlimited open requests. Separate rate limits, concurrent connections and spend ceilings.

  2. Persist uncertain outcomes

    If the provider times out after possible acceptance, preserve the same provider idempotency key and query status where available. If unsupported, delay/reconcile or accept documented duplicate risk; a second provider is not automatically safer.

  3. Process callbacks as evidence

    Verify signatures, deduplicate event IDs and map provider message IDs. Track delivery and complaint/bounce evidence separately where a single monotonic status would discard meaningful later events.

  4. Retry only useful work

    Retry transient quota/network failures with jitter inside the expiry window. Do not retry invalid tokens, permanent hard bounces or suppressed destinations without an explicit re-verification flow.

  5. Read in-app safely

    Inbox entries reference resources whose authorization is checked at read time. Mark-read updates are idempotent; analytics should not infer human reading from provider acceptance.

9. Control amplification and provider recovery

Candidate explains

The dangerous burst is often a campaign expanding to millions of recipients, not the intake rate. I throttle expansion according to downstream capacity and oldest eligible age. Adding workers cannot exceed a provider’s quota or make delayed transactional messages useful again.

DecisionChosen baselineAlternative and trade-off
Channel isolationSeparate queues, worker pools and circuit breakers.One global queue is simpler but an SMS outage can delay all channels.
Audience expansionCheckpointed batches with unique delivery IDs.One task per campaign without progress makes retries expensive and duplication hard to control.
Provider failoverSwitch only before a known unaccepted send, or after reliable reconciliation.Immediate failover after an unknown response may duplicate deliveries.
Push collapseCoalesce replaceable state updates using a business-defined collapse key.Never collapse distinct critical events merely because they share a recipient.
PreferencesCache for expansion efficiency; authoritative final check for strict suppression.Cache-only sending lowers latency but leaves a consent-propagation window.

10. Show operational recovery and its limits

Candidate explains

My dashboard separates accepted-to-eligible, eligible-to-provider and provider-to-receipt latency by category and channel. It includes oldest nonexpired task, unknown outcomes, suppression rate, callback verification failures, bounce/complaint rates and per-provider cost.

FailureDetectionRecovery and remaining limitation
Intake commits but queue publish failsOutbox age increases.Relay republishes; unique delivery keys make duplicate expansion safe.
Worker crashes after provider acceptsLease expires with an unfinished attempt.Reconcile or retry using supported idempotency; otherwise mark uncertainty and document duplicate risk.
Provider outage or quota exhaustionTimeout/429 rate and eligible task age rise.Open circuit, back off and drain at a bounded rate; expire stale tasks instead of sending obsolete alerts.
Unsubscribe races with queued workPreference version changed before authorization.Suppress before send authorization; previously authorized in-flight sends may complete.
Forged or repeated callbackSignature, timestamp or event-ID validation fails.Reject forgery and deduplicate repeats; do not let callbacks authorize new sends.
Destination invalid or reassignedPermanent provider error or verification expiry.Disable the specific token/address binding; require re-verification and avoid retrying stale identities indefinitely.
Campaign overloadTransactional SLO degrades despite aggregate capacity.Reserve transactional capacity, cap expansion and apply tenant fairness; notify operators of delayed campaign completion.

11. Close with an honest delivery contract

Candidate explains

I own durable notification intent, unique recipient/channel expansion and an auditable attempt lifecycle. Providers own the next delivery hop, and users own whether they read the content. I expose those distinctions rather than compressing every stage into a misleading sent flag.

I would finish by testing duplicate business events, a provider timeout after acceptance, unsubscribe during a campaign and out-of-order callbacks. These cases reveal the correctness boundary more effectively than a successful single email demonstration.

Technical references

Primary references explain underlying mechanisms. Workloads and architecture choices above remain proposed interview assumptions.