How to Prepare
Preparation works best as a loop: learn one mechanism, retrieve it from memory, apply it to a new prompt, and review where your explanation failed.
Choose a preparation path
Use the Guided Course if you want concepts, practice and reviews in a deliberate order. Use the Reference Library when a mock exposes a specific gap. Switching paths should not reset your completed lessons or notes: both paths use the same activity identities.
Start with one timed design before building a calendar. Pick a familiar service, spend 35–45 minutes explaining it aloud, and save the diagram. Do not stop to research during this baseline. You are measuring where your reasoning breaks: scope, data ownership, scale, failure recovery or communication.
| Observation in the mock | Next study action | Evidence of improvement |
|---|---|---|
| Named many services but no complete request | Revisit Delivery Framework | Trace one request to durable completion in five minutes |
| Added caches without a freshness contract | Study Caching | Explain a stale-refill race and the selected remedy |
| Said “exactly once” after drawing a queue | Study Multi-step Processes | Handle a crash after commit and before acknowledgement |
| Used numbers without changing a decision | Study Numbers to Know | Identify the capacity threshold behind one choice |
| Ran out of time before failures | Rehearse a shorter baseline | Reach an end-to-end design by the middle of the session |
Build a useful study session
A 90-minute block can contain 15 minutes of retrieval from memory, 25 minutes learning one mechanism, 30 minutes applying it to a new problem, and 20 minutes reviewing the result. These are suggested budgets, not interview rules. The key is producing an explanation before looking at the answer.
For example, learn conditional updates, then apply them to a ticket hold. On another day, apply the same mechanism to an auction bid. If you can repeat the ticket example but cannot explain the auction, you memorized the story more than the guarantee. Ask what state is authoritative, which operation is atomic, and what a timeout leaves uncertain.
A two-week sequence you can adapt
| Days | Main work | Deliverable |
|---|---|---|
| 1–2 | Baseline, Delivery Framework, Networking and API Design | One scoped design with explicit contracts |
| 3–4 | Data Modeling, Indexing, Caching | A read-heavy design with a stale-data failure |
| 5–6 | Sharding, Consistent Hashing, CAP, Numbers | A partition choice and capacity worksheet |
| 7 | Timed Bitly or News Aggregator mock | Saved attempt and evidence-based self-review |
| 8–9 | Contention and multi-step workflows | Ticket or payment state machine with recovery |
| 10–11 | Real-time updates and large files | Reconnect or resumable-upload walkthrough |
| 12 | Long-running work and proximity | A bounded worker design or geospatial search |
| 13 | Unseen full mock | A complete design under a changed constraint |
| 14 | Repair two recurring gaps and retrieve earlier lessons | A concise personal error log |
If you have less time, reduce the number of new problems and preserve the practice-review loop. If you have more time, add spaced transfer exercises rather than rereading the same pages. For senior and staff roles, extend each design with rollout, observability, cost, ownership and the reasoning behind rejected alternatives.
What to write in an error log
Record the prompt, the failed claim, the corrected rule, a counterexample, and a date to retry. “Study Kafka” is too broad. “I assumed committing a consumer offset made an external email send atomic” is specific enough to repair. The next exercise should force that boundary again with a different side effect.
Score five dimensions from 0 to 2: scope, enforceable invariants, complete flows, justified deep dives and clear communication. Zero means absent; one means named; two means demonstrated with a mechanism and failure. A numerical score is a study aid, not a hiring prediction. Keep the actual evidence beside the score so confidence does not substitute for correctness.
Practice with increasing constraints
Start with Bitly, News Aggregator and Rate Limiter for focused designs. Move to Dropbox, Ticketmaster, WhatsApp and Notification System to combine multiple guarantees. Then attempt FB News Feed, Payment System, Google Docs and ChatGPT with strict time limits and unfamiliar follow-ups. These are suggested learning stages, not claims about company interview difficulty.
For every attempt, ask a partner to change one requirement: traffic grows tenfold, a region fails, deletion must propagate promptly, or one tenant produces most of the load. Repair the same design before replacing it wholesale. This reveals whether each component has a clear responsibility.
Before and after a mock
Before starting, prepare a blank canvas, confirm duration and scope, and choose a timer. During the mock, say assumptions aloud and invite correction. At the halfway point, check that every core user action has a path through the architecture. Reserve time for at least one concrete failure and a concise recap.
Afterward, save the original attempt before studying a solution. Compare decisions, not drawing similarity. Choose two changes you can explain precisely and repeat the problem later without notes. Improvement means making fewer unsupported claims and recovering more clearly from follow-ups.
Preparation exercise
You have seven days and repeatedly miss retry behavior. Design three sessions using one concept, two different problems and a delayed retest. State what observable result would convince you that the gap is repaired.
Show answer and explanation
Answer: Study durable acceptance and idempotency, apply them to Notification System, then apply them to Payment System without notes. Two days later, inject a lost provider response into a new workflow. The gap is repaired when you preserve the operation identity, distinguish uncertain from failed, and explain reconciliation without claiming that a queue guarantees exactly-once external effects.
Plan one week of preparation for a candidate who understands APIs but misses concurrency and failure cases.
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.