Spotify Data Lake
Spotify’s RAP account describes an external index that maps lookup keys to files and row locations in existing Parquet data.
Spotify’s RAP account describes an external index that maps lookup keys to files and row locations in existing Parquet data. Cached metadata and targeted range reads avoid a chain of discovery reads. Prepared layouts can reduce bytes or requests further, with tradeoffs for analytical scans. This chapter focuses on online point lookup over a data lake, rather than treating a general data-platform overview as the same case. Source publication: 2026-07. The walkthrough below is an original interview exercise, not an undocumented claim about the company.
The mechanism at a glance
Figure — Lookup key → External index (key lookup); External index → File and row locations (resolve locations); File and row locations → Cached page metadata (page mapping); Cached page metadata → Object range reads (target ranges); Object range reads → Selected records (decode rows); Selected records → Lookup key (answer)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Name the access pattern
For an independent exercise, retrieve one user’s last 100 activity records from a large immutable archive. A full scan is wasteful. Compare copying all records into an online store with keeping a smaller locator index that points into existing files.
2. Make references immutable
A locator must identify a specific file version and a valid row or byte range. If compaction rewrites files, publish new file manifests and corresponding index entries together. Keep old files until active readers and old index generations can no longer reference them.
3. Measure read amplification
Assume a useful record is 200 bytes but the storage page is 1 MB. One successful lookup may still fetch about 5,000 times the payload size. Grouping records by lookup key or storing covering values can reduce that cost, but may make other scans less efficient.
4. Handle missing and changing data
Define whether an absent key means no data or an index that has not caught up. Return a generation or freshness marker. For recently written records, a small hot store can complement the archive, but merging paths needs deduplication and a clear cutover watermark.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
locator(key, dataset_generation, file_version, row_numbers)
manifest(generation, immutable_files, published_at)
cache(file_version, page_offsets)
response(records, generation, freshness)Worked example
Original exercise: a user has records in 30 daily files. Compare 30 independent targeted reads with a weekly layout requiring at most five files. Fewer requests can improve latency and cost, but weekly compaction adds write work and may reduce day-level pruning. Quantify both paths for the query mix instead of assuming one layout is universally better.
Failure walkthrough
Original failure probe: compaction deletes a file while an older index still points to it. A reader receives a missing-object error despite valid data existing elsewhere. Publish immutable generations, delay garbage collection and retry against a new manifest only under a documented consistency policy.
Figure — Read a locator generation → Resolve immutable file versions → Fetch only needed ranges → Decode requested rows → Retire old files after readers finish
Decisions and trade-offs
| Decision | Useful when | Cost to explain |
|---|---|---|
| Adopt the mechanism | The same workload constraint is demonstrated | Validate with your own measurements |
| Keep a simpler design | Your scale and guarantees are already met | Monitor the trigger for changing it |
Check your understanding
Why does a precise index not guarantee a cheap lookup?
Show answer and explanation
Answer: The lookup can identify the correct row while still fetching a large compressed page, many files or many columns. Measure bytes and request count per useful result as well as index lookup time.
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Continue the connection
Study Flink and explain which guarantee from this lesson carries into that topic.
Why does a precise index not guarantee a cheap lookup?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.