Notification System
Design a multi-tenant notification service for transactional alerts and bulk campaigns across email and push. Show the lifecycle from accepted intent to a reconciled outcome.
Save your attempt before continuing. ## Interview scope and guarantees
Accept notification intent, resolve recipients, apply current channel preferences and quiet hours, send through providers, and expose attempt/delivery status. Separate urgent from bulk traffic. Provider acceptance, delivery and user reading are different outcomes.
Capacity worksheet
Assume 10 million recipients receive 2 channels: 20 million delivery jobs. At a combined provider limit of 10,000/s, the theoretical drain time is 2,000 seconds, about 33.3 minutes, before retries. At 200 ms mean provider latency that throughput needs about 2,000 in-flight calls. Increasing workers beyond the provider quota adds contention, not throughput.
Concrete API contract
POST /notifications {requestKey,audienceId,templateId,channels,expiresAt} -> 202 {notificationId}
GET /notifications/{id}/status
POST /provider-events {providerEventId,attemptId,status}
PUT /preferences {channels,quietHours,timeZone}Data model and access paths
intents(id PK,tenant_id,request_key,template_version,state); UNIQUE(tenant_id,request_key)
deliveries(id PK,intent_id,recipient_id,channel,state); UNIQUE(intent_id,recipient_id,channel)
attempts(id PK,delivery_id,provider,key,state)
provider_events(provider,event_id); UNIQUE(provider,event_id)
audience_snapshots(intent_id,snapshot_version,criteria,watermark)
expansion_shards(intent_id,shard,cursor,state); PK(intent_id,shard)
preferences(tenant_id,user_id,channel,version,quiet_hours,time_zone); PK(tenant_id,user_id,channel)
templates(id,version,locale,content_hash); PK(id,version,locale)Evolve a solution and explain each change
Figure — Three architecture decisions for Notification System, including the pressure each introduces.
Step 1: Accept intent
Persist recipient, channel, template version, and deduplication identity. Acceptance is not delivery and can precede user opt-out changes.
Step 2: Apply policy and dispatch
Enforce current preferences and quotas before a provider attempt. Providers can timeout after accepting the message.
Step 3: Reconcile ambiguous attempts
Track attempt identity, provider references, retry classes, and final state. Blind retries can send duplicate irreversible messages.
Responsibility overview
Figure — Connected responsibilities for Notification System. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
A payment event requests email n17. The API commits notification intent and returns accepted. Before dispatch, the worker checks template/version, recipient preferences, and channel policy; transactional and marketing rules can differ. It calls the provider with a stable message identity where supported. The provider accepts but the response is lost. Mark the attempt unknown and query provider status or reconcile a callback instead of immediately treating it as failed. If no idempotency/status mechanism exists, document the at-least-once duplicate risk. A permanent invalid-address error is not retried like a temporary rate limit.
Why these access paths matter
Intent is keyed by producer event, recipient, and purpose. Attempts carry provider, attempt ID, response class, and external reference. Index due attempts by next eligible time; apply quota globally across workers. A delivery webhook is untrusted until verified and deduplicated. Preserve accepted, sent, delivered, bounced, and suppressed as distinct outcomes.
Build the baseline first
- Intent transaction → outbox → audience expansion.
- Delivery queue → preferences and quota → provider.
- Webhook or reconciler → attempt status → reporting.
Evolve the design under load
- Tenant/channel queues → weighted fair dispatch.
- Provider token budgets → bounded worker pools.
- Chunked audience snapshot → checkpointed expansion.
Defend the hardest decision
Freeze the intended audience or define live membership semantics, but recheck opt-outs and legally relevant preferences before sending. Give every recipient/channel delivery a stable ID and reuse its provider idempotency key on ambiguous retries. Quiet-hour delays require timezone-aware due times and message expiry. An expired emergency alert should not be sent hours later merely because the queue recovered.
Failure and recovery analysis
The provider accepts an SMS but the response is lost. Blind failover to a second provider may duplicate the message. Query the first provider by the stable request reference when supported, otherwise wait and reconcile under a documented duplicate-risk policy. A webhook may arrive before the synchronous response and later events may be out of order; use a transition table rather than blindly replacing status.
Security and privacy boundary
Encrypt addresses, verify webhook signatures, authorize tenant templates and audiences, and redact message bodies from routine logs.
Interview follow-ups with reasoning
Question: How do you isolate a large district or tenant?
Show answer and explanation
Answer: Allocate a fair share of dispatch slots and quota with reserved urgent capacity.
Question: How do you recover a partial expansion?
Show answer and explanation
Answer: Store chunk checkpoints and unique delivery keys.
Question: What proves observability?
Show answer and explanation
Answer: Reconcile intent counts to pending, suppressed, accepted, delivered and failed outcomes.
Operate and verify the design
Oldest pending delivery, provider quota usage, uncertain attempts and intent-to-outcome reconciliation.
Lose a provider response and verify the system reconciles before sending through a second provider.
A second scenario to test transfer
A school queues a million notices at 09:00. A parent opts out at 09:10, before their delivery reaches dispatch. The worker checks current preference and suppresses that delivery with an audit reason. Another delivery times out after provider submission; it stays unknown until lookup or callback resolves it. The dashboard should count suppressed and unknown separately from delivered.
Write a design for a million-recipient campaign that includes queue age, opt-outs, and provider timeouts.
Show answer and explanation
Answer: Use durable intent and paged expansion, enforce a provider-wide dispatch budget, recheck live consent, retain attempt identities, and reconcile unknown results. Track acceptance latency, oldest queue age, suppressed counts, provider error rate, and unresolved attempt age separately.
Compare alternatives
| Boundary | Mechanism | Remaining risk |
|---|---|---|
| API to queue | Transactional outbox | Duplicate publication |
| Worker to provider | Stable provider identity | Provider-specific support |
| Consent | Near-dispatch policy check | Race with external submission |
| Callbacks | Authentication + dedupe | Delayed or missing receipts |
A design-changing exercise
The provider returns 202. Can the service report delivered?
Show answer and explanation
Answer: No. Provider acceptance, actual transmission, and delivery confirmation are separate states; expose the evidence available at each boundary.
Design workshop: expand recipients, then enforce current policy
Core scope is email, SMS and auto-call intents, user preferences, provider dispatch and outcome tracking. Exactly-once human receipt is not promised. Choose durable acceptance under 500 ms, a specified campaign delivery window and zero unauthorized new sends under the chosen consent-admission contract. Preference changes after a provider has accepted a message cannot retract it.
Add audience_snapshots(intent_id,snapshot_version,criteria,watermark), expansion_shards(intent_id,shard,cursor,state), preferences(tenant,user,channel,version,quiet_hours,time_zone), and templates(id,version,locale,content_hash). Freeze who belongs to the campaign at a declared snapshot, then evaluate current channel consent/quiet hours near dispatch. Expander inserts deliveries unique(intent,recipient,channel) with its checkpoint in one transaction. A crash repeats a page without creating duplicate logical deliveries.
Figure — Frozen audience, current preference.
A provider quota scheduler combines channel/provider/tenant budgets. At 500 requests/s and 200 ms response latency, Little's Law gives about 100 in-flight API calls, plus headroom bounded by timeout and quota. Auto-calls differ: a two-minute active call duration at 10 new calls/s implies about 1,200 active calls, so connection/concurrent-call quotas may dominate request rate. A 1.5-billion-recipient theoretical school population is not one provider batch; freeze/partition and define a feasible campaign window against purchased limits.
Use weighted tenant rotation, bounded batches and scheduled retry times rather than workers repeatedly hammering a provider. Retry delay=min(cap,base×2^attempt) with jitter; respect Retry-After and expiry. Definitive invalid-address errors are terminal, throttling retries later, timeouts become unknown until the original provider identity/status can be resolved. A failover provider is a new external operation and may duplicate a message if the first outcome remains unknown; do not silently switch on timeout.
Figure — Provider callback transition policy.
Authenticate callback signatures and deduplicate provider/event ID. Separate accepted, delivered, bounced, read where supported, and unknown; provider email accepted is not human read. Persist raw outcome evidence with privacy-aware retention, and derive per-delivery terminal state under provider-specific ordering rules. A reconciler polls old unknown attempts where supported and raises aged ambiguity for operations.
A delivery-count conservation report uses logical recipient/channel deliveries: created = queued+leased+accepted_pending+unknown+delivered+failed+suppressed+expired, with mutually exclusive current categories. Attempt counts are separate; retries must not inflate recipient totals. Reconcile sampled provider records against internal attempt IDs. Track audience lag, preference suppressions, oldest queue age, quota utilization, unknown age and DLQ reason distribution.
Exercise: Email attempt A times out, a failover send B is queued, then A's delivered callback arrives. What happens?
Show answer and explanation
Answer: Reconcile A first and cancel B if it has not crossed its provider admission boundary. If B was already accepted, record possible duplicate delivery honestly. A stable internal delivery identity cannot make two unrelated providers share one external effect.
Stripe webhook guidance illustrates authentication, duplicate and order handling for one provider; email/SMS/call providers require their own documented contracts.