ChatGPT
Design a conversational AI product at the platform level: accepting turns, streaming generated output, managing conversation context, routing inference, and enforcing usage budgets.
Interview scope and guarantees
Accept a conversation turn, stream generated output, retain history and allow cancellation or retry. This is a generic conversational AI service design. Define model availability, context limits, first-token latency and privacy before adding tools or retrieval.
Capacity worksheet
Assume 1,000 requests/s with 10 seconds average generation time: about 10,000 concurrent generations. At 500 output tokens/request the service needs 500,000 output tokens/s. At 2,000 input tokens/request it must also process 2 million input tokens/s; requests/s alone is not an adequate capacity metric.
Concrete API contract
POST /conversations/{id}/turns {clientTurnId,text,model} -> {runId,eventStream}
GET /runs/{id}/events?after=...
POST /runs/{id}/cancel
GET /conversations/{id}Data model and access paths
conversations(id PK,owner_id,retention_policy)
turns(id PK,conversation_id,client_turn_id,request_hash,role,content,state); UNIQUE(conversation_id,client_turn_id)
runs(id PK,turn_id UNIQUE,model_version,context_version,status,current_generation,input_tokens,output_tokens)
run_attempts(run_id,generation,worker_id,lease_until,state); PK(run_id,generation)
stream_events(run_id,sequence,type,payload); PK(run_id,sequence)
tool_actions(run_id,action_id,request_hash,state,provider_ref); PK(run_id,action_id)Evolve a solution and explain each change
Figure — Three architecture decisions for ChatGPT, including the pressure each introduces.
Step 1: Accept a conversation turn
Authenticate access, persist the user turn, and define a generation identity. A retry can start duplicate expensive generations.
Step 2: Stream bounded inference
Admit by token budget and model capacity; emit identified stream events. Client disconnect does not automatically cancel computation.
Step 3: Recover and control context
Store completion state, handle cancellation, and version model/context/tool execution. Generated text and tool requests are not trusted authorization.
Responsibility overview
Figure — Connected responsibilities for ChatGPT. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
A user submits turn t8 and receives generation g8. The inference service emits numbered stream events; the client renders partial text but does not treat it as a completed assistant turn. The connection drops after event 40. Reconnect either resumes retained events for g8 or returns a clearly defined non-resumable status; it must not silently run another billed generation. Cancellation marks the generation and asks the worker to stop, while accounting records work actually consumed. If a tool is invoked, give the tool action its own idempotency and permission boundary. A model-generated instruction cannot grant access to another user’s conversation.
Why these access paths matter
Conversation and turn lookups are scoped to owner/membership. Generations carry turn, model version, context snapshot, status, and token usage; stream events carry sequence and bounded retention. Admission should budget tokens and duration, not merely request count. Context truncation or retrieval selection changes the answer, so record the chosen context boundary for debugging.
Build the baseline first
- Turn API → durable run → inference admission.
- Model worker → streamed token events → client.
- Final message → durable history → usage accounting.
Evolve the design under load
- Token-aware quotas → model-specific queues.
- Prefill/decode scheduling → bounded GPU batches.
- Optional retrieval → authorized evidence → prompt assembly.
Defend the hardest decision
Streaming reduces perceived latency but does not make inference cheap. Separate queue delay, time to first token and generation rate. Long prompts stress prefill; long answers occupy decode capacity and KV cache. Set input/output budgets before admission. Retrying a user turn should not silently create a second billable run; store a stable client-turn identity and define whether retry resumes or starts a new explicit attempt.
Failure and recovery analysis
A browser disconnects halfway through generation. Decide whether to cancel promptly or retain events for resumption; unbounded background runs waste scarce capacity. A tool call with side effects needs its own approval and idempotency boundary. Retrieved documents and model output are untrusted data and cannot grant permission to execute arbitrary actions.
Security and privacy boundary
Authorize conversations and retrieval sources, isolate tool credentials, redact sensitive logs and avoid using private context across users.
Interview follow-ups with reasoning
Question: How do you handle model overload?
Show answer and explanation
Answer: Queue within a deadline, offer an explicit fallback or reject with a retry hint.
Question: How do you evaluate quality?
Show answer and explanation
Answer: Use task-specific test sets, groundedness checks and human review, not latency alone.
Question: How do you enforce deletion?
Show answer and explanation
Answer: Track history, logs, retrieval indexes and backups under a stated retention policy.
Operate and verify the design
Queue delay, time to first token, tokens/s, cancellation waste and retrieval quality.
Disconnect mid-generation, resume or cancel according to policy, and verify no accidental second billable run.
A second scenario to test transfer
A user sends a turn and receives the first 200 tokens before the model worker disconnects. The client has a partial response cursor. The service can resume the same generation if supported, restart as a clearly marked new attempt, or return a recoverable interruption. It should not append an unlabelled second answer that looks like one continuous output.
A generation streams a partial answer, then a tool request times out after possibly succeeding. What should the user and system see?
Show answer and explanation
Answer: Show the generation as interrupted or awaiting tool reconciliation, with a stable turn and action identity. Query or retry the same idempotent tool action when its contract allows; otherwise surface uncertainty for confirmation. Do not present a second unlinked answer or repeat a purchase under a new identity.
Compare alternatives
| Boundary | Guarantee | Key risk |
|---|---|---|
| Turn acceptance | Stable persisted turn | Duplicate submit |
| Token stream | Ordered partial output | Disconnect/resume gap |
| Inference worker | Bounded attempt | Cost and queue overload |
| Tool execution | Authorized auditable action | Ambiguous external effect |
A design-changing exercise
The browser closes. Can the service immediately mark generation failed and retry later?
Show answer and explanation
Answer: Only under a stated lifecycle policy. The worker may still be generating; retain its identity, coordinate cancellation, and prevent an automatic duplicate generation.
Design workshop: one turn, bounded inference, typed stream
Core scope is conversation turns, streamed generation, cancellation and authorized tools. Model training is excluded; RAG is a separately scoped extension. Choose time-to-first-token under two seconds for admitted short prompts, explicit context-window limits and a resumable event horizon. GPU backlog admission must be honest: a queued response is not an already running stream.
Unique(conversation,clientTurnId) maps to one logical run with request fingerprint. Persist model/context version, current attempt generation and typed stream_events(run,sequence,type,payload). Types include started, token_delta, tool_requested, tool_result, completed, interrupted and canceled. GET /runs/id/events?after=N replays retained events; expired history returns reset_required with durable current run status. A retry with changed text/model under the same turn identity returns 409.
Figure — Resume the same run after token 40.
Cancellation and completion serialize on run state. The first valid terminal transition wins under a declared rule; a cancel request marks cancel_requested, revokes publication ownership if cancellation wins, signals the worker and records consumed usage. Already delivered tokens remain visible as partial output. A stale worker cannot append another completed answer after the run is canceled. Distinguish one logical run from potentially repeated costly inference attempts; charge once according to the selected billing policy and retain attempt evidence for reconciliation.
Context assembly applies access checks before retrieving conversation/documents, counts tokens using the selected model tokenizer and pins exact input evidence. For an 8,192-token context limit, reserve 1,024 output tokens and 512 system/tool tokens, leaving 6,656 for user/history/retrieval. Truncate or summarize old turns under a visible policy; do not silently cut critical tool state. A retained summary carries source/version so replay/debugging can reconstruct what was sent.
Figure — Illustrative decoder KV cache budget.
The formula assumes the specified grouped-query attention architecture and uncompressed two-byte KV values; other models/attention schemes differ. Prefill processes prompt tokens and can block latency-sensitive decode if scheduled carelessly. Continuous batching admits/finishes sequences per decoding step while a token/KV-memory budget limits total active work. Chunked prefill and bounded prompt/output limits protect decode fairness. Routing assigns model/version-specific replicas; worker queues are isolated by token and GPU memory budget, not merely requests/s.
Measure prompt-token throughput, generated tokens/s, active KV bytes, queue wait, first-token and inter-token latency separately. Prefix reuse is allowed only with model/context/version-safe cache keys and tenant/privacy policy; a cache hit must not expose another user's content. A replica failure marks the attempt interrupted. Resume retained emitted events where possible, but do not pretend an arbitrary lost GPU execution state can resume generation from exactly the same internal state.
For a RAG extension: authorized documents→versioned chunking→retrieval candidates→reranking→bounded context with citations. Evaluate answer grounding and retrieval recall using versioned test sets; retrieval content is untrusted and cannot authorize tools. Tools enforce user/account scopes and operation identities outside model instructions. Ambiguous external effects query original action evidence rather than issue an unrelated purchase/send.
Exercise: Tool A may have succeeded before timeout and the inference worker crashes. May a restarted run simply invoke another tool action B?
Show answer and explanation
Answer: No. Reconcile A with its provider key/status under the original permission boundary. Surface uncertainty if it cannot be resolved. A new run/attempt does not grant a new side effect or justify duplicate billing.
vLLM paged attention is a primary implementation reference for paged KV storage; the numerical model and product state machine above are explicit teaching assumptions.
Technical references
A generation streams a partial answer, then a tool request times out after possibly succeeding. What should the user and system see?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.