Learning pathsA
GUIDED PRACTICE

ChatGPT

Design a conversational AI product at the platform level: accepting turns, streaming generated output, managing conversation context, routing inference, and enforcing usage budgets.

Interview scope and guarantees

Accept a conversation turn, stream generated output, retain history and allow cancellation or retry. This is a generic conversational AI service design. Define model availability, context limits, first-token latency and privacy before adding tools or retrieval.

Capacity worksheet

Assume 1,000 requests/s with 10 seconds average generation time: about 10,000 concurrent generations. At 500 output tokens/request the service needs 500,000 output tokens/s. At 2,000 input tokens/request it must also process 2 million input tokens/s; requests/s alone is not an adequate capacity metric.

Concrete API contract

Contract / pseudocode
POST /conversations/{id}/turns {clientTurnId,text,model} -> {runId,eventStream}
GET /runs/{id}/events?after=...
POST /runs/{id}/cancel
GET /conversations/{id}

Data model and access paths

Contract / pseudocode
conversations(id PK,owner_id,retention_policy)
turns(id PK,conversation_id,client_turn_id,request_hash,role,content,state); UNIQUE(conversation_id,client_turn_id)
runs(id PK,turn_id UNIQUE,model_version,context_version,status,current_generation,input_tokens,output_tokens)
run_attempts(run_id,generation,worker_id,lease_until,state); PK(run_id,generation)
stream_events(run_id,sequence,type,payload); PK(run_id,sequence)
tool_actions(run_id,action_id,request_hash,state,provider_ref); PK(run_id,action_id)

Evolve a solution and explain each change

Three architecture decisions for ChatGPT, including the pressure each introduces.
Scroll to inspect the diagram, or open it at full size.

Figure — Three architecture decisions for ChatGPT, including the pressure each introduces.

Step 1: Accept a conversation turn

Authenticate access, persist the user turn, and define a generation identity. A retry can start duplicate expensive generations.

Step 2: Stream bounded inference

Admit by token budget and model capacity; emit identified stream events. Client disconnect does not automatically cancel computation.

Step 3: Recover and control context

Store completion state, handle cancellation, and version model/context/tool execution. Generated text and tool requests are not trusted authorization.

Responsibility overview

Connected responsibilities for ChatGPT. Trace the authoritative and derived paths separately.
Scroll to inspect the diagram, or open it at full size.

Figure — Connected responsibilities for ChatGPT. Trace the authoritative and derived paths separately.

Worked end-to-end scenario

A user submits turn t8 and receives generation g8. The inference service emits numbered stream events; the client renders partial text but does not treat it as a completed assistant turn. The connection drops after event 40. Reconnect either resumes retained events for g8 or returns a clearly defined non-resumable status; it must not silently run another billed generation. Cancellation marks the generation and asks the worker to stop, while accounting records work actually consumed. If a tool is invoked, give the tool action its own idempotency and permission boundary. A model-generated instruction cannot grant access to another user’s conversation.

Why these access paths matter

Conversation and turn lookups are scoped to owner/membership. Generations carry turn, model version, context snapshot, status, and token usage; stream events carry sequence and bounded retention. Admission should budget tokens and duration, not merely request count. Context truncation or retrieval selection changes the answer, so record the chosen context boundary for debugging.

Build the baseline first

ChatGPT: baseline request paths.
Scroll to inspect the diagram, or open it at full size.
  1. Turn API → durable run → inference admission.
  2. Model worker → streamed token events → client.
  3. Final message → durable history → usage accounting.

Evolve the design under load

ChatGPT: additional scaling and recovery paths.
Scroll to inspect the diagram, or open it at full size.
  1. Token-aware quotas → model-specific queues.
  2. Prefill/decode scheduling → bounded GPU batches.
  3. Optional retrieval → authorized evidence → prompt assembly.

Defend the hardest decision

Streaming reduces perceived latency but does not make inference cheap. Separate queue delay, time to first token and generation rate. Long prompts stress prefill; long answers occupy decode capacity and KV cache. Set input/output budgets before admission. Retrying a user turn should not silently create a second billable run; store a stable client-turn identity and define whether retry resumes or starts a new explicit attempt.

Failure and recovery analysis

A browser disconnects halfway through generation. Decide whether to cancel promptly or retain events for resumption; unbounded background runs waste scarce capacity. A tool call with side effects needs its own approval and idempotency boundary. Retrieved documents and model output are untrusted data and cannot grant permission to execute arbitrary actions.

Security and privacy boundary

Authorize conversations and retrieval sources, isolate tool credentials, redact sensitive logs and avoid using private context across users.

Interview follow-ups with reasoning

Question: How do you handle model overload?

Show answer and explanation

Answer: Queue within a deadline, offer an explicit fallback or reject with a retry hint.

Question: How do you evaluate quality?

Show answer and explanation

Answer: Use task-specific test sets, groundedness checks and human review, not latency alone.

Question: How do you enforce deletion?

Show answer and explanation

Answer: Track history, logs, retrieval indexes and backups under a stated retention policy.

Operate and verify the design

Queue delay, time to first token, tokens/s, cancellation waste and retrieval quality.

Disconnect mid-generation, resume or cancel according to policy, and verify no accidental second billable run.

A second scenario to test transfer

A user sends a turn and receives the first 200 tokens before the model worker disconnects. The client has a partial response cursor. The service can resume the same generation if supported, restart as a clearly marked new attempt, or return a recoverable interruption. It should not append an unlabelled second answer that looks like one continuous output.

A generation streams a partial answer, then a tool request times out after possibly succeeding. What should the user and system see?

Show answer and explanation

Answer: Show the generation as interrupted or awaiting tool reconciliation, with a stable turn and action identity. Query or retry the same idempotent tool action when its contract allows; otherwise surface uncertainty for confirmation. Do not present a second unlinked answer or repeat a purchase under a new identity.

Compare alternatives

BoundaryGuaranteeKey risk
Turn acceptanceStable persisted turnDuplicate submit
Token streamOrdered partial outputDisconnect/resume gap
Inference workerBounded attemptCost and queue overload
Tool executionAuthorized auditable actionAmbiguous external effect

A design-changing exercise

The browser closes. Can the service immediately mark generation failed and retry later?

Show answer and explanation

Answer: Only under a stated lifecycle policy. The worker may still be generating; retain its identity, coordinate cancellation, and prevent an automatic duplicate generation.

Design workshop: one turn, bounded inference, typed stream

Core scope is conversation turns, streamed generation, cancellation and authorized tools. Model training is excluded; RAG is a separately scoped extension. Choose time-to-first-token under two seconds for admitted short prompts, explicit context-window limits and a resumable event horizon. GPU backlog admission must be honest: a queued response is not an already running stream.

Unique(conversation,clientTurnId) maps to one logical run with request fingerprint. Persist model/context version, current attempt generation and typed stream_events(run,sequence,type,payload). Types include started, token_delta, tool_requested, tool_result, completed, interrupted and canceled. GET /runs/id/events?after=N replays retained events; expired history returns reset_required with durable current run status. A retry with changed text/model under the same turn identity returns 409.

Resume the same run after token 40. Client to Turn/event authority: Submit client turn T8; get run R8; Turn/event authority to Inference pool: Admit attempt generation 3 with pinned context; Inference pool to Turn/event authority: Append numbered token deltas 1..40; Client to Turn/event authority: Reconnect after 40, not submit a new billed turn; Inference pool to Tool service: Tool action A under separate permission/idempotency check; Inference pool to Turn/event authority: Conditional completed event; reject revoked attempt generation
Scroll to inspect the diagram, or open it at full size.

Figure — Resume the same run after token 40.

Cancellation and completion serialize on run state. The first valid terminal transition wins under a declared rule; a cancel request marks cancel_requested, revokes publication ownership if cancellation wins, signals the worker and records consumed usage. Already delivered tokens remain visible as partial output. A stale worker cannot append another completed answer after the run is canceled. Distinguish one logical run from potentially repeated costly inference attempts; charge once according to the selected billing policy and retain attempt evidence for reconciliation.

Context assembly applies access checks before retrieving conversation/documents, counts tokens using the selected model tokenizer and pins exact input evidence. For an 8,192-token context limit, reserve 1,024 output tokens and 512 system/tool tokens, leaving 6,656 for user/history/retrieval. Truncate or summarize old turns under a visible policy; do not silently cut critical tool state. A retained summary carries source/version so replay/debugging can reconstruct what was sent.

Illustrative decoder KV cache budget. 32 layers, 8 KV heads, head dimension 128 / 2 for K/V × 32 × 8 × 128 × 2-byte values / 131,072 bytes/token; 8,192 tokens for one sequence / 128 KiB × 8,192 / 1 GiB KV before allocator overhead; 32 such sequences / 32 × 1 GiB / 32 GiB KV plus model weights/workspace; More heads/layers or longer contexts / Linear in those dimensions for this model / Request count alone is insufficient admission metric
Scroll to inspect the diagram, or open it at full size.

Figure — Illustrative decoder KV cache budget.

The formula assumes the specified grouped-query attention architecture and uncompressed two-byte KV values; other models/attention schemes differ. Prefill processes prompt tokens and can block latency-sensitive decode if scheduled carelessly. Continuous batching admits/finishes sequences per decoding step while a token/KV-memory budget limits total active work. Chunked prefill and bounded prompt/output limits protect decode fairness. Routing assigns model/version-specific replicas; worker queues are isolated by token and GPU memory budget, not merely requests/s.

Measure prompt-token throughput, generated tokens/s, active KV bytes, queue wait, first-token and inter-token latency separately. Prefix reuse is allowed only with model/context/version-safe cache keys and tenant/privacy policy; a cache hit must not expose another user's content. A replica failure marks the attempt interrupted. Resume retained emitted events where possible, but do not pretend an arbitrary lost GPU execution state can resume generation from exactly the same internal state.

For a RAG extension: authorized documents→versioned chunking→retrieval candidates→reranking→bounded context with citations. Evaluate answer grounding and retrieval recall using versioned test sets; retrieval content is untrusted and cannot authorize tools. Tools enforce user/account scopes and operation identities outside model instructions. Ambiguous external effects query original action evidence rather than issue an unrelated purchase/send.

Exercise: Tool A may have succeeded before timeout and the inference worker crashes. May a restarted run simply invoke another tool action B?

Show answer and explanation

Answer: No. Reconcile A with its provider key/status under the original permission boundary. Surface uncertainty if it cannot be resolved. A new run/attempt does not grant a new side effect or justify duplicate billing.

vLLM paged attention is a primary implementation reference for paged KV storage; the numerical model and product state machine above are explicit teaching assumptions.

Technical references

Kafka processing and external-sink boundaries.

PostgreSQL transaction isolation and concurrent updates.

12:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

A generation streams a partial answer, then a tool request times out after possibly succeeding. What should the user and system see?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.