Learning pathsA
In a Hurry

Delivery Framework

A complete interview framework with a worked notification design.

Outcome: take an ambiguous prompt to a working design, explain its most important limits, and leave time for a useful conversation about trade-offs.

You can know a lot about databases and still struggle in a design interview. The hard part is deciding what to explain next. You draw a cache, remember that caches need invalidation, start describing replication, and suddenly twenty minutes have passed without a complete request path.

A framework gives you a way back. At any moment, you should be able to answer: which user action am I designing, which requirement does this component satisfy, and what remains unresolved?

We will use a notification service throughout this lesson. A school administrator creates a notice, chooses an audience, and asks the system to send it through each recipient's preferred channel. This example is fictional; its traffic and latency targets are assumptions for the exercise. It is useful because the happy path is simple, while retries and external providers force us to be precise.

A practical 45-minute agenda

An original notebook-style roadmap showing scope, state, contracts, baseline, pressure tests, and recap.
Scroll to inspect the diagram, or open it at full size.

Figure 1. Make one complete design early enough to spend the remaining time testing it. Optional data-flow reasoning belongs inside the architecture discussion when a pipeline needs it.

StageExample budgetWhat should exist when you move on?
Agree on scope5 minutesA short feature list, explicit exclusions, and measurable quality targets
Identify important state3 minutesNamed entities, ownership, and one or two invariants
Define contracts5 minutesCore operations, responses, and relevant retry behavior
Build a working baseline12 minutesOne end-to-end path for each agreed user action
Test the design under pressure15 minutesTwo justified deep dives, with architecture changes where necessary
Recap5 minutesRequirements met, key trade-offs, and remaining risks

This adds up to 45 minutes, but it is a suggested budget, not an interview rule. If the interviewer asks about consistency early, follow that discussion. The purpose of the agenda is to keep you aware of the whole problem while adapting to the conversation.

For a shorter interview, reduce scope before rushing. For a senior interview, use extra depth to defend correctness and operations, not just to add components.

1. Agree on what the system must do

Start with people and actions. “Build a notification platform” could mean marketing campaigns, urgent alerts, transactional receipts, or all three. Those products have different deadlines, preference rules, and failure consequences.

A useful opening is:

I’ll focus on an administrator sending a notice to an authorized audience, recipients receiving it through their allowed channel, and the administrator checking delivery progress. Do we need recurring schedules or replies in this design?

You have proposed a workable scope and left the interviewer room to correct it.

Functional requirements

For this exercise, agree on three capabilities:

  1. An authorized administrator creates and submits a notice for a selected audience.
  2. The service dispatches a delivery through an eligible channel after applying the relevant preference and consent policy.
  3. The administrator can inspect progress and individual failures.

Exclude template editing, billing, replies, and recurring schedules from the first version. Excluding something does not mean it is unimportant in production. It means you have chosen what this conversation will establish.

Non-functional requirements

“Fast, scalable, and reliable” will not tell us whether the design works. Attach each quality to an operation or failure.

For our example, assume:

  • The submission API should normally acknowledge a valid request within 300 ms at the 95th percentile. This is acceptance latency, not the time until every recipient receives a message.
  • Once the API acknowledges acceptance, the notice intent must survive an API process restart. Database durability and recovery policies still need to be specified for broader infrastructure failures.
  • A burst of one million email deliveries should not exceed the provider's agreed request budget.
  • One tenant's campaign must not consume every dispatch slot.
  • Delivery visibility must distinguish queued work, provider acceptance, confirmed delivery, and a final failure.

Consent requirements depend on the product and jurisdiction; for the design exercise, assume an explicit policy that is checked before dispatch. Do not invent a universal rule for every communication channel.

State the invariant

An invariant is a condition that must remain true despite concurrency and retries. For this system:

A delivery attempt belongs to one logical delivery, and retries must not silently create a different business operation.

This does not promise that no recipient can ever receive a duplicate. If the provider lacks idempotency support and returns an ambiguous timeout, a duplicate window can remain. Naming that limitation is better than writing “exactly once” next to a queue.

2. Estimate a number when it can change a decision

You do not need a page of arithmetic before drawing anything. Estimate when a quantity can rule out an approach, establish a bound, or select a partitioning strategy.

Suppose our hypothetical email provider allows 500 submission requests per second, and this exercise uses one request per recipient. One million messages have a theoretical dispatch lower bound of:

Contract / pseudocode
1,000,000 requests / 500 requests per second
= 2,000 seconds
≈ 33 minutes 20 seconds

That is not an end-to-end delivery guarantee. Retries, other tenants, provider throttling, and the provider's own delivery pipeline can make it longer.

Now we can explain an architectural choice: accepting a campaign synchronously and waiting for all messages is incompatible with our API response target. We need asynchronous processing and honest progress reporting. Adding ten thousand workers will not remove the provider's limit.

A second useful estimate is in-flight work. If an individual provider call takes an average of 200 ms, sustaining 500 calls per second requires roughly 100 concurrent calls under stable conditions:

Contract / pseudocode
average in-flight work ≈ throughput × average time
                       ≈ 500/s × 0.2s
                       ≈ 100 calls

This is a starting estimate, not a configured safe limit. Real concurrency must account for latency variation, connection capacity, errors, and quotas. It gives us a reason to discuss bounded dispatch rather than unlimited parallelism.

3. Name the important state

List the entities that make the requirements concrete. You can refine fields while tracing requests.

EntityWhat it representsImportant fields for this discussion
NoticeThe administrator's intenttenantId, noticeId, audienceRef, contentRef, status
DeliveryOne logical recipient/channel operationdeliveryId, noticeId, recipientId, channel, status
AttemptA particular submission attemptattemptId, deliveryId, providerRef, outcome, nextRetryAt
PreferenceCurrent channel eligibility under policyrecipientId, channel, policyVersion, enabled
Outbox eventA committed fact awaiting publicationeventId, noticeId, type, publishedAt
Provider receiptAn external status observationproviderEventId, providerRef, observedStatus, timestamp

Ask where each piece of truth lives. The database owns the accepted notice and delivery state. A queue transports work. A dashboard counter is a derived view. These roles are not interchangeable.

Use stable logical identifiers. If a worker invents a new delivery ID on every retry, it becomes much harder to recognize that two attempts belong to one intended message.

Avoid modelling every profile field or every communication preference at this stage. Add fields when a query, invariant, or failure path needs them.

4. Define the user-visible contract

Show enough API detail to make your design testable. REST is a reasonable example choice here; RPC and GraphQL can also be appropriate depending on the interfaces and workloads. Protocol names alone do not establish performance.

http
POST /v1/notices
Authorization: Bearer <token>
Idempotency-Key: <client-generated-operation-key>

{
  "audienceId": "audience_83",
  "templateVersion": "v12",
  "parameters": { "eventDate": "2026-10-20" }
}
json
{
  "noticeId": "notice_218",
  "status": "accepted",
  "statusUrl": "/v1/notices/notice_218"
}

Return 202 Accepted after the required durable acceptance transaction succeeds. Do not return it merely because the request arrived in process memory.

Derive the caller's identity from authenticated context, then authorize access to the tenant and audience. A valid token does not automatically permit access to any audience identifier supplied in the request.

Define the idempotency key's scope, retention period, and payload fingerprint. Reusing a key with the same operation should return the recorded result. Reusing it with a different payload should be rejected as a conflict. The idempotency record and accepted notice must be committed atomically, otherwise concurrent submissions can still create duplicates.

The read side could be:

http
GET /v1/notices/{noticeId}
GET /v1/notices/{noticeId}/deliveries?status=failed&limit=50&cursor=...

Explain cursor pagination if the audience can be large. A response that returns millions of delivery records is neither a useful UI contract nor a safe default API.

5. Build one complete baseline

For a pipeline, write down the work before choosing every component:

  1. Authenticate and authorize the submission.
  2. Commit notice intent and a publication record.
  3. Expand the audience in bounded batches.
  4. Check eligibility and create logical deliveries.
  5. Dispatch attempts under tenant and provider budgets.
  6. Record results and reconcile provider receipts.
A notebook-style architecture separating durable acceptance from asynchronous dispatch and provider receipts.
Scroll to inspect the diagram, or open it at full size.

Figure 2. The acceptance transaction establishes durable intent. Later boundaries establish different facts. A provider accepting a request is not the same as the recipient receiving it.

For the first diagram, use a notice API, PostgreSQL, an outbox publisher, a durable work queue, dispatch workers, and the provider. A receipt handler records asynchronous delivery outcomes.

Trace a request out loud. The API writes the notice, idempotency result, and outbox entry in one transaction. A publisher reads committed outbox work and sends it to the queue. It may publish twice if it crashes after publishing but before recording success, so consumers must tolerate duplicate events.

The dispatcher acquires work, checks the applicable eligibility policy, and submits through the correct channel under configured limits. It records the provider reference when known. A provider callback updates the delivery view using a deduplicated receipt event.

Do not treat the outbox as magic. It closes the database-to-message publication gap under the publisher's retry process, but does not by itself make the provider effect exactly once. We still need a strategy at that external boundary.

At this point, the system covers all three user actions. We can now spend the remaining time on the hardest requirements without leaving the main path unfinished.

6. Choose deep dives from the requirements

Deep dive A: a successful provider call looks like a timeout

A notebook-style sequence showing provider acceptance followed by a lost response, leaving the worker uncertain.
Scroll to inspect the diagram, or open it at full size.

Figure 3. A timeout tells the worker that the result is unknown. It does not prove that the provider rejected the operation.

Suppose the provider accepted a message but the response was lost. Retrying with a new logical identity can send a second message.

If supported, reuse the same provider idempotency key and reconcile against the original reference. If a status lookup is available, use it to resolve uncertainty. If neither is available, document the product decision: retry and accept a possible duplicate, or delay/escalate and accept a possible omission. The design cannot manufacture a guarantee the external interface does not provide.

Explain the recovery mechanism, its delay, and who can observe unresolved attempts. An “unknown” status is often more truthful and operationally useful than declaring every timeout a failure.

Deep dive B: a burst exceeds downstream capacity

A queue lets producers and consumers work at different rates. It does not make a sustained rate mismatch disappear.

Bound campaign expansion, queue age, worker concurrency, database connections, and provider request rate. Apply tenant fairness so a large campaign cannot monopolize every slot. Introduce admission control or defer new campaigns when the system cannot meet its stated completion policy.

Monitor the age of the oldest pending delivery as well as the queue length. A short queue containing stuck work can be worse than a long queue that drains predictably. Track provider throttle responses, retry volume, connection-pool wait, and completion time by tenant and channel.

A useful interviewer follow-up is: “Which limit is authoritative if we double the worker fleet?” Your answer should point to a shared budget or coordination mechanism; an independent 500/s limiter in each worker would multiply the permitted global rate.

Deep dive C: preferences change while a campaign waits

Separate audience membership from permission to dispatch. A campaign can preserve an audience snapshot for auditing while rechecking current suppression or consent at send time under the product's rules.

Record why a delivery was skipped and which policy version produced the decision. If preferences are cached, define the allowable staleness and the path for urgent changes. A cache TTL is not a sufficient answer when the requirement says a revocation must take effect immediately.

There is still a concurrency boundary between the check and the external call. If strict revocation ordering is required, discuss how dispatch and preference changes are serialized or version-checked and define the point after which an already-submitted message cannot be recalled.

You probably cannot explore all three dives fully in one interview. Choose the two most relevant to the prompt and adapt to the interviewer's questions.

7. Explain what you will operate

Reserve a short moment for signals and recovery. This makes the architecture more than a diagram of successful requests.

Useful signals include API acceptance latency, oldest unprocessed outbox age, queue age, provider throttling, unknown-result count, dispatch latency, and delivery outcome rates. Use noticeId, deliveryId, and attemptId to join traces and logs without placing sensitive message content in every log entry.

Alerts should connect to a response. Growing outbox age suggests publication is stuck. A rising unknown-result count suggests reconciliation needs attention. Provider throttling while the queue grows can require adjusting dispatch policy rather than adding workers.

For security, enforce tenant isolation on reads and writes, restrict provider credentials to dispatch components, minimize sensitive payload retention, and protect callbacks with the provider's supported verification mechanism. These choices should follow the data and trust boundaries you already drew.

8. End with a concise recap

A useful close might be:

We acknowledge a notice after committing durable intent. The outbox and queue move that intent to bounded workers, while provider references and receipts give us delivery visibility. We protect provider capacity and tenant fairness independently of worker count. The main unresolved guarantee is duplicate behavior when an external provider offers neither idempotency nor a reliable status lookup.

That summary shows what the system does, why the major components exist, and where its limits remain. Leave room for the interviewer to ask about one of those limits.

Try it yourself

Take 12 minutes. Design only the opening and first architecture for a large file-export service.

  • Pick three user actions and two measurable quality requirements.
  • Write the entities and the submission/status APIs.
  • Mark when the API can safely return success or acceptance.
  • Explain what happens if a worker dies after uploading a file but before updating the job record.
  • State one bottleneck you would investigate before adding workers.

Then close your notes and explain the design for three minutes. Notice where you rely on a technology name instead of a guarantee.

Self-review rubric

A strong answer separates acceptance from completion, identifies durable job state, makes retries safe with a stable job/attempt identity, and defines ownership of the final result. It also explains cleanup for orphaned files and bounds concurrent work. Several storage and queue choices can satisfy these requirements; the reasoning matters more than a particular vendor.

Check your understanding

1. The API responds before committing the notice. What remains possible?

A. A retry cannot happen. B. The accepted intent can disappear if the process fails. C. The provider must send a duplicate.

Show answer and explanation

Answer: B. The acknowledgment promised acceptance before the system established the required durable state. A is false because clients can retry after ambiguity; C is not implied because the provider may never have received anything.

2. A provider call times out. Which conclusion is justified?

A. The provider did not accept the message. B. The provider delivered the message. C. The result is unknown to the caller.

Show answer and explanation

Answer: C. A timeout alone establishes neither acceptance nor rejection. Reconcile or retry using the provider's supported duplicate-prevention contract.

3. Workers can submit 2,000 messages/s, but the shared provider budget is 500/s. What is the dispatch ceiling before other constraints?

A. 500/s. B. 2,000/s. C. Unlimited with a queue.

Show answer and explanation

Answer: A. Worker capacity does not override an external budget. A queue absorbs temporary excess arrival rate; sustained excess still grows the backlog.

4. A user disables a channel after audience expansion. What should the design specify?

A. That audience membership always overrides preferences. B. When eligibility is checked and which policy applies to queued deliveries. C. That a cache automatically invalidates itself.

Show answer and explanation

Answer: B. A is a product rule we have not established. C assumes an invalidation guarantee that a cache does not automatically supply.

5. What is the best reason to add a component during the interview?

A. It appears in many architecture diagrams. B. It makes the diagram look more advanced. C. It satisfies a stated requirement or addresses a demonstrated limitation.

Show answer and explanation

Answer: C. Popularity and diagram complexity are not evidence that a component belongs in this design.

Primary documentation

PostgreSQL transactions

AWS transactional outbox

Practice the conversation, not just the diagram

A framework becomes useful when it changes what you say under pressure. Here is an original miniature exchange for a notification service.

Interviewer promptCandidate reasoning
“Send to every parent at once.”Clarify whether at once means accepted immediately or delivered within a deadline. Provider quotas determine delivery time.
“Why add a queue?”It buffers the accepted burst and permits bounded retries; it does not increase provider capacity.
“Can two workers send the same message?”Yes under repeated delivery unless the effect is protected. Use stable delivery identity and provider idempotency where supported.
“What if the provider times out?”Treat the result as uncertain, retain the same request key, and reconcile before unsafe failover.
“How do we know nothing was lost?”Reconcile accepted intents to expanded, suppressed, pending and terminal delivery counts, and alert on old unresolved work.

Decide which deep dive earns time

A deep dive should resolve a tension between requirements. Low latency versus current data suggests caching and consistency. High throughput versus one-owner correctness suggests contention and partitioning. Fast acceptance versus slow external work suggests queues, leases and recovery. Draw the simplest baseline first so the deep dive changes a visible boundary rather than adding an unrelated service.

For a senior-level answer, describe how you would migrate safely. Introduce a new projection beside the old one, backfill it, compare sampled results, shift traffic gradually and keep a rollback point. Specify what happens to writes during cutover. “Use a feature flag” is incomplete unless both sides remain compatible and accepted work cannot disappear.

A final two-minute recap

Name the three core actions your design supports. Restate the strongest invariant and the operation that enforces it. Give one capacity calculation that affected a decision. Explain one remaining tradeoff, one failure that is handled, and one risk requiring measurement. Do not spend the recap listing every database name.

Transfer exercise

Apply the agenda to a file-sharing service. At minute 20, the interviewer asks for offline concurrent edits. What do you change?

Show answer and explanation

Answer: Clarify whether the requirement means whole-file conflict copies or character-level collaborative merging. Preserve the upload and version-storage baseline. Add version checks and conflict behavior for whole-file edits; only introduce an operation-based collaboration model if the product needs it. Explicitly rebudget the remaining time around the new correctness problem.

Your study notes