Scaling Reads
Read scaling is a sequence of decisions about cost and freshness.
Read scaling is a sequence of decisions about cost and freshness. Begin by finding what each request actually does. A slow query, repeated identical work, and a geographically distant database require different fixes. Adding replicas before understanding the query can copy the same expensive work to more machines without solving its cause.
Learning goals
Measure expensive reads; optimize queries; indexes; replicas; cache layers; materialized views; read-your-writes; overload protection.
The mechanism at a glance
Figure — Read classifier → Fresh owner read (read your writes); Read classifier → Cache (repeated public data); Cache → Replica (bounded miss); Read classifier → Projection (precomputed view); Authoritative writes → Replica (replicate); Authoritative writes → Projection (events)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Remove unnecessary work
Measure the expensive query, selected columns, number of round trips, and access plan. Fix N+1 lookups and add an index that matches the filter and ordering. Limit result sizes. A well-shaped query can delay a distributed redesign and simplifies every later layer.
2. Separate fresh reads from tolerant reads
Read replicas can serve browsing and reporting when bounded lag is acceptable. A user reading immediately after a successful write may need leader routing or a session consistency mechanism. Replica lag is time-varying, so a fixed delay is not a proof that every replica has caught up. State what happens when lag exceeds the product budget.
3. Cache repeated answers
Cache at the browser, edge, application, or object level according to who may share the answer. Include authorization context in keys where needed. Coalesce misses and bound origin concurrency so losing the cache does not destroy the primary store. Choose TTL and invalidation based on a freshness contract.
4. Precompute expensive projections
A materialized catalog or feed can move joins and ranking off the request path. Keep the authoritative data and the projection distinct. Record a watermark or version, make consumers replayable, and plan a backfill that does not mix incompatible schemas. Projection lag should be visible to both operators and, when relevant, users.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
Read classes
A: account change confirmation -> authoritative owner
B: public catalog -> cache, then replica
C: daily report -> materialized projection
All classes: bounded deadline and bounded origin concurrencyWorked example
A catalog gets 10,000 reads/s and 100 writes/s. First index category and sort order. Then cache popular category pages for 30 seconds if the product permits it. Route the owner’s post-edit confirmation to the leader so their change is visible immediately. Add replicas for cold reads once measurement shows the leader’s read load is still the bottleneck. Each step has a named reason.
Failure walkthrough
A replica falls minutes behind while remaining reachable. A load balancer that checks only TCP health will continue routing stale reads to it. Monitor replay lag and remove unsuitable replicas from the fresh-read pool. Avoid sending all failed replica traffic to the leader without a limit; degrade optional pages or apply admission control instead.
Figure — User commits profile edit → Replica has older version → Session carries write context → Route confirmation to owner → Later reads use caught-up replica
Decisions and trade-offs
| Mechanism | Reduces | Freshness implication |
|---|---|---|
| Index | Rows and pages examined | Same transactional source |
| Replica | Primary read load | May lag |
| Cache | Repeated computation | TTL/invalidation contract |
| Projection | Online join work | Asynchronous update lag |
Check your understanding
A profile update succeeds, but the next read shows the old name. Which read should be rerouted?
Show answer and explanation
Answer: The post-write confirmation read should use the authoritative owner or a replica proven to have applied the relevant write. Sending every read to the leader may be unnecessary, while waiting an arbitrary fixed interval does not prove catch-up.
Transfer to a new scenario
Move a frequently read catalog through increasingly costly scaling options, stating when each becomes justified.
A replica returns old data after a successful write. Which request should go to the leader?
Continue the connection
Study PostgreSQL and explain which guarantee from this lesson carries into that topic.
Choose the cheapest useful read path
Begin with the actual query plan and the shape of the response. An index, smaller result, or precomputed count can eliminate work before distributing it. Add read replicas when the workload tolerates their freshness behavior; route a read that must observe a recent write to an appropriate authority or use a version-aware mechanism.
A cache is worthwhile when requests reuse results. A million unique queries may have almost no hits and still consume memory. Materialized views move work from read time to update time, which helps repeated complex reads but introduces lag and rebuild responsibilities. CDN caching works particularly well for immutable public bytes; personalized responses need careful cache keys and authorization.
Estimate the amplification at each layer. A page that makes twenty internal calls can produce a large backend load even at modest page QPS. Batch compatible reads, bound fanout and avoid turning one slow dependency into a full-page timeout.
Figure — A decision worksheet for Scaling Reads: read the mechanism and its guarantee together.
Operational sketch
profile read -> indexed primary query
repeat-heavy profile read -> cache-aside
large dashboard join -> maintained projection
recent write requiring freshness -> authoritative readA tempting mistake
Replicas do not reduce write work and can amplify storage cost. A cache outage must have a bounded fallback; otherwise the normal high hit rate hides an origin that cannot survive misses.
Transfer exercise
When should you avoid adding a cache?
Show answer and explanation
Answer: When reuse is low, freshness requirements make invalidation disproportionately complex, or a straightforward query/index change already meets the target.
A profile update succeeds, but the next read shows the old name. Which read should be rerouted?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.