Learning pathsA
Core Concepts

Caching

A cache stores a reusable answer closer to the caller or cheaper than the source.

A cache stores a reusable answer closer to the caller or cheaper than the source. Its central design question is how wrong that answer is allowed to be. Public article metadata may tolerate a minute of staleness; the final inventory decision may not. State the freshness contract before selecting TTLs, invalidation mechanisms, or a cache product.

Learning goals

Cache-aside; write-through; invalidation; TTL; staleness contracts; stampedes; jitter; request coalescing; negative caching; hot keys; outage protection.

The mechanism at a glance

Reader → Cache lookup (lookup); Cache lookup → Return cached value (hit); Cache lookup → Bounded loader (miss); Bounded loader → Source database (read); Source database → Versioned fill (value + version); Versioned fill → Cache lookup (populate)
Scroll to inspect the diagram, or open it at full size.

Figure — Reader → Cache lookup (lookup); Cache lookup → Return cached value (hit); Cache lookup → Bounded loader (miss); Bounded loader → Source database (read); Source database → Versioned fill (value + version); Versioned fill → Cache lookup (populate)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. Trace cache-aside explicitly

On a hit, return the cached value if its scope and freshness satisfy the request. On a miss, read the authoritative store and populate the cache with a bounded lifetime. Include tenant, authorization-sensitive scope, and representation version in the key when needed. Cache “not found” carefully and briefly so newly created data does not remain invisible for too long.

2. Handle concurrent fills

If 10,000 clients miss the same key at once, a cache can amplify load rather than reduce it. Coalesce fills for that key, use jittered expirations, and limit total origin concurrency. A lock used only to prevent a stampede is not a business correctness lock. It needs a safe timeout and failure behavior rather than indefinite waiting.

3. Explain invalidation races

Deleting a key after a write is not enough in every interleaving. Reader A can fetch old database state, writer B can commit and invalidate, then A can refill the cache with the old value. Versioned entries, coordinated invalidation, or a short bounded TTL can address different freshness contracts. Choose a mechanism proportional to the consequence of stale data.

4. Design the outage path

If the cache disappears, unrestricted fallback can overload the database. Apply admission control, cap concurrent misses, shed optional work, and serve explicitly permitted stale values where useful. A high hit rate is not sufficient evidence of health; a few expensive misses or one hot key may dominate origin load and tail latency.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
Illustrative load calculation
Read traffic = 20,000 requests/s
Hit ratio = 95%
Expected misses = 1,000 requests/s
If cache fails: potential origin load = 20,000 requests/s
Origin protection must be sized for its safe capacity.

Worked example

A public article changes its title. A one-minute stale title is acceptable, so use cache-aside with jittered TTL and a best-effort invalidation event. A checkout availability display can also be cached, but the purchase operation must revalidate inventory at its owner. This separation makes the interface responsive without letting a cached “in stock” value authorize an oversell.

Failure walkthrough

During a synchronized TTL expiry, all web workers try to rebuild a popular homepage. Add one in-flight refresh per key and a global fill budget. If stale-while-revalidate is allowed, return the old page during a bounded refresh window. If freshness is mandatory, reject or wait within a deadline rather than silently serve stale data. Monitor miss concurrency, eviction rate, fill latency, and stale age.

Reader fetches old version → Writer commits new version → Writer invalidates cache → Old reader attempts refill → Version check or TTL bounds stale data
Scroll to inspect the diagram, or open it at full size.

Figure — Reader fetches old version → Writer commits new version → Writer invalidates cache → Old reader attempts refill → Version check or TTL bounds stale data

Decisions and trade-offs

StrategyBenefitRisk
Cache-asideSimple and demand-drivenMiss latency and stale fills
Write-throughCache updated on write pathWrite latency and dual-system failures
Short TTLBounds ordinary stale lifetimeMore misses; not an atomicity guarantee

Check your understanding

A key expires during a database slowdown. Which matters more: adding cache clients or bounding cache-miss work?

Show answer and explanation

Answer: Bound miss work and coalesce refreshes. More clients do not increase database capacity. A protected origin can recover; an unlimited retry and refill wave can keep it saturated.

Transfer to a new scenario

Cache public article metadata while keeping the final inventory decision authoritative.

A cache miss wave reaches an already saturated database. Which protection actually bounds load?

Continue the connection

Study Scaling Reads and explain which guarantee from this lesson carries into that topic.

Define freshness before choosing a policy

For each value, state who owns the truth, how stale the value may be, and whether stale data can cause harm. A product description may tolerate a minute of delay. A seat reservation must still perform its conflict check at the authoritative store. Caching a displayed availability count does not make that count safe for purchase decisions.

Use a key that includes every dimension affecting the response: tenant, entity, relevant query parameters and possibly schema version. Caching an authenticated response under only a URL can leak one user’s data to another. Large values and high-cardinality query combinations can exhaust memory even with a good hit rate.

Compare read and write policies

PolicyFlowPrincipal risk
Cache-asideApplication reads cache, then store on a missStale refill races and stampedes
Read-throughCache loader fetches a missLoader becomes a critical dependency
Write-throughWrite path updates durable state and cache under a defined protocolDual-write failure still needs handling
Write-behindCache acknowledges before durable persistenceLost acknowledged data if the buffer fails
Invalidate after writeCommit source, then delete cached valueInvalidation can be delayed or lost

A policy name is not a consistency proof. “Write-through” across two independent systems still needs an order, retry behavior and recovery story. If acknowledged data must survive a cache loss, an ordinary volatile cache cannot be the only committed copy.

Walk through the stale-refill race

Reader R misses the cache and reads database version 7. Before R fills the cache, writer W commits version 8 and invalidates the key. R then inserts version 7. The invalidation succeeded, yet stale data returned. A short TTL limits the stale interval but does not prevent the race.

Two concurrent readers and writers demonstrating a stale refill after invalidation.
Scroll to inspect the diagram, or open it at full size.

Possible remedies include version-aware cache writes with a retained invalidation version, a coherent loader/write protocol, or bypassing the cache for operations requiring current state. A delayed second delete reduces some races but is not a universal proof under arbitrary delays. State which residual staleness your choice permits.

Prevent a cache outage becoming a database outage

With 100,000 reads/s and a 99% hit rate, the database normally sees around 1,000 reads/s. Losing the cache can multiply that load by 100. Adding application replicas makes the failure worse if every instance immediately falls back to the database.

Coalesce concurrent misses for the same key, cap origin concurrency, apply per-tenant admission and use stale-while-revalidate only where the freshness contract permits it. Add TTL jitter so a bulk-loaded set does not expire simultaneously. Negative caching can protect nonexistent keys, but its TTL must not hide newly created records for too long.

Separate eviction from expiry

Expiry implements a lifetime rule. Eviction makes room under memory pressure, potentially before a TTL expires. LRU favors recency; LFU favors repeated use; approximate policies reduce bookkeeping cost. Measure bytes as well as entry count. A few huge values can evict many useful small ones.

A hot key may need replication or local caching; adding hash partitions will not distribute one key’s requests. Local caches reduce network traffic but multiply invalidation targets. Track hit rate by traffic class and value size, origin requests, fill latency, evictions, stale responses and refresh failures.

Exercise: choose a safe fallback

A recommendation cache is unavailable while an inventory cache is stale. Both normally reduce database load. Should the services use the same fallback?

Show answer and explanation

Answer: No. Recommendations may use a bounded stale result or a simpler fallback. Inventory purchase must still validate authoritative availability and may reject or queue under overload. Apply separate budgets so a recommendation cache outage cannot consume every database connection needed for checkout.

Your study notes