Learning pathsA
GUIDED PRACTICE

Spotify Data Lake

Spotify’s RAP account describes an external index that maps lookup keys to files and row locations in existing Parquet data.

Spotify’s RAP account describes an external index that maps lookup keys to files and row locations in existing Parquet data. Cached metadata and targeted range reads avoid a chain of discovery reads. Prepared layouts can reduce bytes or requests further, with tradeoffs for analytical scans. This chapter focuses on online point lookup over a data lake, rather than treating a general data-platform overview as the same case. Source publication: 2026-07. The walkthrough below is an original interview exercise, not an undocumented claim about the company.

The mechanism at a glance

Lookup key → External index (key lookup); External index → File and row locations (resolve locations); File and row locations → Cached page metadata (page mapping); Cached page metadata → Object range reads (target ranges); Object range reads → Selected records (decode rows); Selected records → Lookup key (answer)
Scroll to inspect the diagram, or open it at full size.

Figure — Lookup key → External index (key lookup); External index → File and row locations (resolve locations); File and row locations → Cached page metadata (page mapping); Cached page metadata → Object range reads (target ranges); Object range reads → Selected records (decode rows); Selected records → Lookup key (answer)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. Name the access pattern

For an independent exercise, retrieve one user’s last 100 activity records from a large immutable archive. A full scan is wasteful. Compare copying all records into an online store with keeping a smaller locator index that points into existing files.

2. Make references immutable

A locator must identify a specific file version and a valid row or byte range. If compaction rewrites files, publish new file manifests and corresponding index entries together. Keep old files until active readers and old index generations can no longer reference them.

3. Measure read amplification

Assume a useful record is 200 bytes but the storage page is 1 MB. One successful lookup may still fetch about 5,000 times the payload size. Grouping records by lookup key or storing covering values can reduce that cost, but may make other scans less efficient.

4. Handle missing and changing data

Define whether an absent key means no data or an index that has not caught up. Return a generation or freshness marker. For recently written records, a small hot store can complement the archive, but merging paths needs deduplication and a clear cutover watermark.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
locator(key, dataset_generation, file_version, row_numbers)
manifest(generation, immutable_files, published_at)
cache(file_version, page_offsets)
response(records, generation, freshness)

Worked example

Original exercise: a user has records in 30 daily files. Compare 30 independent targeted reads with a weekly layout requiring at most five files. Fewer requests can improve latency and cost, but weekly compaction adds write work and may reduce day-level pruning. Quantify both paths for the query mix instead of assuming one layout is universally better.

Failure walkthrough

Original failure probe: compaction deletes a file while an older index still points to it. A reader receives a missing-object error despite valid data existing elsewhere. Publish immutable generations, delay garbage collection and retry against a new manifest only under a documented consistency policy.

Read a locator generation → Resolve immutable file versions → Fetch only needed ranges → Decode requested rows → Retire old files after readers finish
Scroll to inspect the diagram, or open it at full size.

Figure — Read a locator generation → Resolve immutable file versions → Fetch only needed ranges → Decode requested rows → Retire old files after readers finish

Decisions and trade-offs

DecisionUseful whenCost to explain
Adopt the mechanismThe same workload constraint is demonstratedValidate with your own measurements
Keep a simpler designYour scale and guarantees are already metMonitor the trigger for changing it

Check your understanding

Why does a precise index not guarantee a cheap lookup?

Show answer and explanation

Answer: The lookup can identify the correct row while still fetching a large compressed page, many files or many columns. Measure bytes and request count per useful result as well as index lookup time.

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Continue the connection

Study Flink and explain which guarantee from this lesson carries into that topic.

6:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

Why does a precise index not guarantee a cheap lookup?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.