Strava
Design activity recording and fitness discovery: mobile uploads, route samples, summaries, segment matching, and social feeds.
Interview scope and guarantees
Record activities, upload GPS samples, calculate summaries, compare segment efforts and control activity privacy. Offline upload and resumable transfer matter. Raw observations and derived leaderboards must be separately versioned.
Capacity worksheet
Assume 1 million activities/day with 3,600 samples each: 3.6 billion samples/day. At 32 compressed bytes/sample that is about 115 GB/day. Processing 100 candidate segments per activity means 100 million geometric comparisons/day; use spatial pruning before exact matching.
Concrete API contract
POST /activities/uploads {deviceId,clientActivityId} -> {activityId,uploadUrl}
POST /activities/{id}/complete {checksum}
GET /activities/{id}
GET /segments/{id}/leaderboard?cursor=...Data model and access paths
activities(id PK,athlete_id,start_time,privacy,raw_key,version)
samples(activity_id,sequence,time,lat,lon,altitude)
efforts(segment_id,athlete_id,activity_id,duration,algorithm_version)
segment_geometry(segment_id,shape,bounding_cells)Evolve a solution and explain each change
Figure — Three architecture decisions for Strava, including the pressure each introduces.
Step 1: Save an activity
Persist the route and original samples with owner visibility. Publicly projecting precise coordinates can expose sensitive locations.
Step 2: Process asynchronously
Validate samples and derive summaries and segment efforts from a pinned route version. GPS noise and retries can generate misleading or duplicate results.
Step 3: Publish privacy-safe views
Apply privacy transformations before public indexing and map rendering. Hiding a marker alone can leave coordinates in API responses.
Responsibility overview
Figure — Connected responsibilities for Strava. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
A runner uploads activity a8 and chooses followers-only visibility with a hidden start zone. Keep the original samples protected. A worker validates timestamps, computes distance and segment attempts, and writes results keyed by activity version and algorithm version. A replay replaces the same result rather than appending duplicate leaderboard attempts. The public or follower map uses transformed samples and cannot expose the hidden coordinates in its JSON. If visibility changes later, invalidate feed/map projections and enforce access when serving. Segment leaderboards need an explicit rule for private activities and anomalous GPS or speed data.
Why these access paths matter
Activity lookup is keyed by ID and owner; sample storage uses activity and sequence/time. Derived efforts use activity, segment, and derivation version. Spatial indexes contain only the coordinates allowed for their audience. Bounding-box filtering needs geometric refinement, and segment matching needs tolerance rather than exact point equality.
Build the baseline first
- Device upload → immutable raw track → activity record.
- Processing job → cleaned samples → distance and pace.
- Spatial candidates → exact segment match → efforts.
Evolve the design under load
- Partitioned activity workers → versioned calculations.
- Per-segment ranking → bounded leaderboard cache.
- Privacy events → projection removal and recheck.
Defend the hardest decision
GPS contains jitter, gaps and impossible jumps. Preserve raw input and record cleaning parameters so derived distances can be reproduced. Retrieve candidate segments from intersecting cells, then match ordered geometry and direction; passing near one point does not prove completion. Segment scores require an explicit timing rule and a tie-breaker.
Failure and recovery analysis
A privacy change arrives while a leaderboard job is running. Commit only if the activity version and eligibility still match, or ensure the serving layer checks current eligibility. Reprocessing must replace the previous contribution for the same activity and algorithm generation instead of adding another effort. Device retries deduplicate by athlete and client activity ID.
Security and privacy boundary
Hide sensitive start/end locations according to policy and never expose private tracks through public segment responses.
Interview follow-ups with reasoning
Question: How do you handle long offline tracks?
Show answer and explanation
Answer: Multipart upload and sequence checks avoid memory-bound ingestion.
Question: How do you correct an algorithm bug?
Show answer and explanation
Answer: Replay raw tracks into a new derived generation and compare results.
Question: How do you prevent cheating?
Show answer and explanation
Answer: Flag implausible speed and device anomalies without treating every GPS error as fraud.
Operate and verify the design
Processing backlog, rejected GPS anomalies, projection version and privacy cleanup delay.
Reprocess an activity after a privacy change and verify it does not return to a public leaderboard.
A second scenario to test transfer
An athlete records a two-hour ride offline. The phone uploads in chunks 1–20; it loses connectivity after chunk 14 and retries chunks 12–14. Stable chunk identities prevent duplicate samples. Finalize only after required ranges and checksums pass. Processing can run after the upload without holding an API worker for the transfer.
How should the upload API handle a chunk retry, and how should a later processor version replace an earlier summary?
Show answer and explanation
Answer: Make chunk identity include activity and sequence range, verify checksum, and treat a duplicate identical chunk as success while rejecting conflicting bytes. Derive a new summary under a versioned computation and conditionally publish only if its source version is current. Preserve the raw source for reproducible recomputation.
Compare alternatives
| Stage | Guarantee | Risk |
|---|---|---|
| Chunk upload | Resumable ordered samples | Overlap and missing ranges |
| Finalize | Complete raw activity identity | Partial visibility |
| Derived metrics | Versioned computation | Old worker overwrites |
| Privacy projection | Safe public route | Endpoint leakage |
A design-changing exercise
The map hides the first kilometre but the API returns every GPS point. Is the privacy zone working?
Show answer and explanation
Answer: No. Apply the transformation at the data-serving boundary, including downloadable and indexed representations, not only in the visual map.
Design workshop: route geometry, privacy, and versioned efforts
Core flows are chunk upload, activity summary, segment matching and leaderboard. Live navigation is excluded. Assume summaries available within one minute under admitted load, resumable uploads and no raw private endpoints exposed by public APIs. GPS inaccuracies must be marked, not silently converted into performance records.
Store upload_chunks(activity,sequence_range,checksum,object_key) with unique range identity, an expected-range manifest, activity(source_version,state,owner), and derived_outputs(activity,processor_version,source_version,state). Overlapping ranges with differing bytes return conflict. Finalization verifies complete ordered ranges/checksums, publishes an immutable raw track and emits a versioned processing job.
For a segment in a local teaching coordinate system from start (0,0) to finish (1,000,0), use a 20-meter endpoint/corridor tolerance and forward traversal. Samples at (-10,0,t=0), (10,0,t=2), (990,0,t=100) and (1010,0,t=102) cross start at t=1 and finish at t=101 by linear interpolation, giving 100 seconds. A reversed trace is not this directed segment. Sparse samples that could have taken a different road must not be accepted solely from endpoint proximity; require intermediate corridor/order evidence and a maximum sample-gap policy.
Figure — A directed segment's interpolated crossings.
Use a spatial index to find candidate segments near the route, then exact geometry and ordered crossing checks. Persist efforts(activity,segment,source_version,algorithm_version,start_index,end_index,duration,validity). Select a clear leaderboard policy: one best eligible effort per athlete/segment, then duration ascending, completion time and athlete ID as ties. Deleting an activity recomputes that athlete's best remaining effort, not merely decrements a global score.
Figure — Recompute without stale workers winning.
A privacy zone is a protected area around home/work, applied server-side before public projection and map rendering. Clip route portions inside the zone and avoid returning raw samples, hidden bounds or API endpoints that reveal the original route. A segment effort revealing a private start/end or traversal needs the same policy; hiding map markers alone is insufficient. Bound caching by permission/privacy version and regenerate projections after zone changes. Previously downloaded raw routes cannot be remotely erased by a UI change.
At one sample/s for two hours, an activity has 7,200 points. At an illustrative 32 bytes/point, that is 230,400 bytes raw before attributes and object overhead. One million such activities/day produce about 230 GB/day. Processing cost depends on route points × candidate segments, so cap candidate retrieval, simplify only with known error bounds and preserve raw evidence for reproducible computation.
Exercise: Algorithm 4 corrects a duration from 90 to 100 seconds while an algorithm-3 worker finishes late. Which result becomes current?
Show answer and explanation
Answer: The output matching current source and processor version wins through conditional publication. Retain the old result as historical evidence if desired; update per-athlete best effort and affected segment ranking from the new authoritative projection.
Technical references
How should the upload API handle a chunk retry, and how should a later processor version replace an earlier summary?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.