Practice: Notification System
Design a multi-tenant notification service for transactional alerts and bulk campaigns across email and push. Show the lifecycle from accepted intent to a reconciled outcome.
Design a multi-tenant notification service for transactional alerts and bulk campaigns across email and push. Show the lifecycle from accepted intent to a reconciled outcome.
Write your own design before opening hints or the solution review. The numbers below are exercise assumptions, not claims about any company’s production traffic. You may challenge an assumption, but record the replacement and explain which decision changes.
Scenario and constraints
- Largest campaign: one million recipients.
- Email provider budget: 500 submissions/s globally.
- Average provider call time: 200 ms; quotas are illustrative.
- Consent may change after expansion; urgent alerts outrank bulk work.
Your deliverables
- Model campaign, delivery, attempt, preference, and receipt identities.
- Calculate the dispatch lower bound and expected in-flight provider work.
- Draw durable acceptance, paged expansion, fair dispatch, and receipts.
- Define queued, suppressed, submitted, delivered, failed, and unknown states.
Reason about this sequence
Figure — Intent accepted durably → Recipient opts out → Worker checks current consent → Suppress before provider call → Record auditable outcome
For every step, annotate what is durable, what the caller knows, and which identity survives a retry. Identify the point where two concurrent actors could disagree. Do not assume a timeout means failure or a cache value grants ownership.
Interviewer follow-ups
- The database commits but queue publication fails.
- The provider accepts a send but its response is lost.
- A recipient opts out while the campaign waits in the queue.
Answer each follow-up using the same design first. If it breaks, change the smallest boundary that repairs the invariant and explain the new cost. Show whether the change adds latency, storage, coordination, or operational work.
Staged hints
Hint 1 — reveal
Answer: Hint 1 — Use separate identities for logical delivery and each provider attempt. Queue durability does not tell you whether a timed-out provider request succeeded.
Hint 2 — reveal
Answer: Hint 2 — A queued audience snapshot is not permanent permission to send. Recheck current consent and suppression rules close to the provider call, and document the unavoidable boundary between that check and an external send. Apply quiet hours, channel preferences, and per-tenant fairness before consuming scarce provider capacity.
Hint 3 — reveal
Answer: Hint 3 — The database commits a campaign but the process crashes before publishing to the broker. Without an outbox, accepted work can disappear. A transactionally stored outbox event can be published later. Publication may repeat, so recipient expansion and dispatch still need stable identities. The outbox closes one gap; it does not create exactly-once external delivery.
Evidence-based self-review
- Acceptance cannot lose committed intent at the broker boundary.
- Global provider capacity is enforced across all workers.
- Consent and suppression are checked under a stated dispatch policy.
- Unknown provider outcomes are reconciled with stable identities.
Score each item 0 if absent, 1 if named without an enforceable mechanism, or 2 if the mechanism and a failure are explained. Record evidence from your own diagram beside the score. Then choose one weak decision, revise it, and repeat the relevant follow-up. This rubric is a learning tool, not a hiring forecast.
Figure — The worked notification design separates acceptance, dispatch policy, provider attempts, and reconciliation.
Worked resolution
There is no separate review activity for this prompt. Open the explanation after saving your attempt.
Show answer and explanation
Answer: A school queues a million notices at 09:00. A parent opts out at 09:10, before their delivery reaches dispatch. The worker checks current preference and suppresses that delivery with an audit reason. Another delivery times out after provider submission; it stays unknown until lookup or callback resolves it. The dashboard should count suppressed and unknown separately from delivered. Use durable intent and paged expansion, enforce a provider-wide dispatch budget, recheck live consent, retain attempt identities, and reconcile unknown results. Track acceptance latency, oldest queue age, suppressed counts, provider error rate, and unresolved attempt age separately.
Show answer and explanation
Answer: Capacity check — one million / 500 per second = 2,000 seconds before retries and overhead. At 200 ms average provider latency, 500/s implies about 100 in-flight operations in a stable system. Global quotas still constrain every worker pool together.
Reveal the full worked solution
Save your attempt before continuing. ## Interview scope and guarantees
Accept notification intent, resolve recipients, apply current channel preferences and quiet hours, send through providers, and expose attempt/delivery status. Separate urgent from bulk traffic. Provider acceptance, delivery and user reading are different outcomes.
Capacity worksheet
Assume 10 million recipients receive 2 channels: 20 million delivery jobs. At a combined provider limit of 10,000/s, the theoretical drain time is 2,000 seconds, about 33.3 minutes, before retries. At 200 ms mean provider latency that throughput needs about 2,000 in-flight calls. Increasing workers beyond the provider quota adds contention, not throughput.
Concrete API contract
POST /notifications {requestKey,audienceId,templateId,channels,expiresAt} -> 202 {notificationId}
GET /notifications/{id}/status
POST /provider-events {providerEventId,attemptId,status}
PUT /preferences {channels,quietHours,timeZone}Data model and access paths
intents(id PK,tenant_id,request_key,template_version,state); UNIQUE(tenant_id,request_key)
deliveries(id PK,intent_id,recipient_id,channel,state); UNIQUE(intent_id,recipient_id,channel)
attempts(id PK,delivery_id,provider,key,state)
provider_events(provider,event_id); UNIQUE(provider,event_id)
audience_snapshots(intent_id,snapshot_version,criteria,watermark)
expansion_shards(intent_id,shard,cursor,state); PK(intent_id,shard)
preferences(tenant_id,user_id,channel,version,quiet_hours,time_zone); PK(tenant_id,user_id,channel)
templates(id,version,locale,content_hash); PK(id,version,locale)Evolve a solution and explain each change
Figure — Three architecture decisions for Notification System, including the pressure each introduces.
Step 1: Accept intent
Persist recipient, channel, template version, and deduplication identity. Acceptance is not delivery and can precede user opt-out changes.
Step 2: Apply policy and dispatch
Enforce current preferences and quotas before a provider attempt. Providers can timeout after accepting the message.
Step 3: Reconcile ambiguous attempts
Track attempt identity, provider references, retry classes, and final state. Blind retries can send duplicate irreversible messages.
Responsibility overview
Figure — Connected responsibilities for Notification System. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
A payment event requests email n17. The API commits notification intent and returns accepted. Before dispatch, the worker checks template/version, recipient preferences, and channel policy; transactional and marketing rules can differ. It calls the provider with a stable message identity where supported. The provider accepts but the response is lost. Mark the attempt unknown and query provider status or reconcile a callback instead of immediately treating it as failed. If no idempotency/status mechanism exists, document the at-least-once duplicate risk. A permanent invalid-address error is not retried like a temporary rate limit.
Why these access paths matter
Intent is keyed by producer event, recipient, and purpose. Attempts carry provider, attempt ID, response class, and external reference. Index due attempts by next eligible time; apply quota globally across workers. A delivery webhook is untrusted until verified and deduplicated. Preserve accepted, sent, delivered, bounced, and suppressed as distinct outcomes.
Build the baseline first
- Intent transaction → outbox → audience expansion.
- Delivery queue → preferences and quota → provider.
- Webhook or reconciler → attempt status → reporting.
Evolve the design under load
- Tenant/channel queues → weighted fair dispatch.
- Provider token budgets → bounded worker pools.
- Chunked audience snapshot → checkpointed expansion.
Defend the hardest decision
Freeze the intended audience or define live membership semantics, but recheck opt-outs and legally relevant preferences before sending. Give every recipient/channel delivery a stable ID and reuse its provider idempotency key on ambiguous retries. Quiet-hour delays require timezone-aware due times and message expiry. An expired emergency alert should not be sent hours later merely because the queue recovered.
Failure and recovery analysis
The provider accepts an SMS but the response is lost. Blind failover to a second provider may duplicate the message. Query the first provider by the stable request reference when supported, otherwise wait and reconcile under a documented duplicate-risk policy. A webhook may arrive before the synchronous response and later events may be out of order; use a transition table rather than blindly replacing status.
Security and privacy boundary
Encrypt addresses, verify webhook signatures, authorize tenant templates and audiences, and redact message bodies from routine logs.
Interview follow-ups with reasoning
Question: How do you isolate a large district or tenant?
Show answer and explanation
Answer: Allocate a fair share of dispatch slots and quota with reserved urgent capacity.
Question: How do you recover a partial expansion?
Show answer and explanation
Answer: Store chunk checkpoints and unique delivery keys.
Question: What proves observability?
Show answer and explanation
Answer: Reconcile intent counts to pending, suppressed, accepted, delivered and failed outcomes.
Operate and verify the design
Oldest pending delivery, provider quota usage, uncertain attempts and intent-to-outcome reconciliation.
Lose a provider response and verify the system reconciles before sending through a second provider.
A second scenario to test transfer
A school queues a million notices at 09:00. A parent opts out at 09:10, before their delivery reaches dispatch. The worker checks current preference and suppresses that delivery with an audit reason. Another delivery times out after provider submission; it stays unknown until lookup or callback resolves it. The dashboard should count suppressed and unknown separately from delivered.
Write a design for a million-recipient campaign that includes queue age, opt-outs, and provider timeouts.
Show answer and explanation
Answer: Use durable intent and paged expansion, enforce a provider-wide dispatch budget, recheck live consent, retain attempt identities, and reconcile unknown results. Track acceptance latency, oldest queue age, suppressed counts, provider error rate, and unresolved attempt age separately.
Compare alternatives
| Boundary | Mechanism | Remaining risk |
|---|---|---|
| API to queue | Transactional outbox | Duplicate publication |
| Worker to provider | Stable provider identity | Provider-specific support |
| Consent | Near-dispatch policy check | Race with external submission |
| Callbacks | Authentication + dedupe | Delayed or missing receipts |
A design-changing exercise
The provider returns 202. Can the service report delivered?
Show answer and explanation
Answer: No. Provider acceptance, actual transmission, and delivery confirmation are separate states; expose the evidence available at each boundary.
Design workshop: expand recipients, then enforce current policy
Core scope is email, SMS and auto-call intents, user preferences, provider dispatch and outcome tracking. Exactly-once human receipt is not promised. Choose durable acceptance under 500 ms, a specified campaign delivery window and zero unauthorized new sends under the chosen consent-admission contract. Preference changes after a provider has accepted a message cannot retract it.
Add audience_snapshots(intent_id,snapshot_version,criteria,watermark), expansion_shards(intent_id,shard,cursor,state), preferences(tenant,user,channel,version,quiet_hours,time_zone), and templates(id,version,locale,content_hash). Freeze who belongs to the campaign at a declared snapshot, then evaluate current channel consent/quiet hours near dispatch. Expander inserts deliveries unique(intent,recipient,channel) with its checkpoint in one transaction. A crash repeats a page without creating duplicate logical deliveries.
Figure — Frozen audience, current preference.
A provider quota scheduler combines channel/provider/tenant budgets. At 500 requests/s and 200 ms response latency, Little's Law gives about 100 in-flight API calls, plus headroom bounded by timeout and quota. Auto-calls differ: a two-minute active call duration at 10 new calls/s implies about 1,200 active calls, so connection/concurrent-call quotas may dominate request rate. A 1.5-billion-recipient theoretical school population is not one provider batch; freeze/partition and define a feasible campaign window against purchased limits.
Use weighted tenant rotation, bounded batches and scheduled retry times rather than workers repeatedly hammering a provider. Retry delay=min(cap,base×2^attempt) with jitter; respect Retry-After and expiry. Definitive invalid-address errors are terminal, throttling retries later, timeouts become unknown until the original provider identity/status can be resolved. A failover provider is a new external operation and may duplicate a message if the first outcome remains unknown; do not silently switch on timeout.
Figure — Provider callback transition policy.
Authenticate callback signatures and deduplicate provider/event ID. Separate accepted, delivered, bounced, read where supported, and unknown; provider email accepted is not human read. Persist raw outcome evidence with privacy-aware retention, and derive per-delivery terminal state under provider-specific ordering rules. A reconciler polls old unknown attempts where supported and raises aged ambiguity for operations.
A delivery-count conservation report uses logical recipient/channel deliveries: created = queued+leased+accepted_pending+unknown+delivered+failed+suppressed+expired, with mutually exclusive current categories. Attempt counts are separate; retries must not inflate recipient totals. Reconcile sampled provider records against internal attempt IDs. Track audience lag, preference suppressions, oldest queue age, quota utilization, unknown age and DLQ reason distribution.
Exercise: Email attempt A times out, a failover send B is queued, then A's delivered callback arrives. What happens?
Show answer and explanation
Answer: Reconcile A first and cancel B if it has not crossed its provider admission boundary. If B was already accepted, record possible duplicate delivery honestly. A stable internal delivery identity cannot make two unrelated providers share one external effect.
Stripe webhook guidance illustrates authentication, duplicate and order handling for one provider; email/SMS/call providers require their own documented contracts.
Technical references
Write a design for a million-recipient campaign that includes queue age, opt-outs, and provider timeouts.
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.