News Aggregator
Design a personalized list of articles collected from publishers.
Interview scope and guarantees
Ingest feeds, normalize stories, group duplicates, rank recent items and return a paginated feed. Start with public syndicated sources. Define freshness, attribution and what happens when publishers correct or retract a story.
Capacity worksheet
Assume 100,000 sources polled every 15 minutes: about 111 polls/s before retries. One million articles/day at 10 KB is 10 GB/day raw. Ten million users reading 5 times/day generate 579 reads/s average; cache popular category feeds rather than reranking the full corpus for every request.
Concrete API contract
POST /sources {feedUrl,category}
GET /feed?category=...&cursor=...
GET /stories/{id}
POST /stories/{id}/feedback {action}Data model and access paths
sources(id PK,url,etag,next_poll_at,last_success)
articles(id PK,source_id,canonical_url,published_at,content_hash,version)
clusters(cluster_id,article_id,similarity_version)
feed_generation(category,generation,ordered_story_ids)Evolve a solution and explain each change
Figure — Three architecture decisions for News Aggregator, including the pressure each introduces.
Step 1: Collect source articles
Fetch feeds under source limits and preserve provenance. The same story can arrive with different URLs and titles.
Step 2: Normalize and index
Assign article identity, cluster related stories, and build read projections. Similarity is not proof that two reports are interchangeable.
Step 3: Serve personalized pages
Rank bounded candidates with freshness and source diversity, using stable cursors. Re-ranking a changing list can duplicate pages or hide updates.
Responsibility overview
Figure — Connected responsibilities for News Aggregator. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
Two feeds publish the same syndicated article under different URLs. URL normalization catches only some duplicates; content and publisher metadata can help identify a common article. A different reporter’s story about the same event belongs in the same cluster but should keep its own identity and attribution. Fetchers persist source versions before indexing. A corrected headline updates the article version, and the projection rejects an older replay. A reader page uses a ranking snapshot or cursor contract so an inserted breaking story does not shift offsets unpredictably. If the projection is unavailable, a simpler recent-articles view can be a defined degradation.
Why these access paths matter
Store source URL, publisher, fetched version, timestamps, and licensing/provenance attributes. Index articles by publication time and category; cluster membership is a revisable projection. User preferences belong in a separate access-controlled record. Keep fetching cadence per source and distinguish ingestion lag from ranking freshness.
Build the baseline first
- Polite poller → raw feed → normalized articles.
- Deduplication → story clusters → ranker.
- Published feed generation → cache → reader.
Evolve the design under load
- Source queues → retry isolation → bounded fetchers.
- Incremental ranking → category snapshots.
- Feedback stream → bounded personalization features.
Defend the hardest decision
Canonical URLs and exact hashes catch obvious duplicates; semantic clustering can group related reporting but may incorrectly merge different events. Retain original article identities and attribution even when presenting a story cluster. Publication time supplied by a feed may be wrong or updated, so use ingestion time for pipeline monitoring and a documented ranking policy for freshness.
Failure and recovery analysis
One publisher begins returning malformed or enormous XML. Apply parser limits, timeouts and quarantine to that source without blocking ingestion globally. If an article is retracted, propagate a versioned tombstone through clusters and cached feeds. A ranked snapshot gives stable pagination while newer generations remain available on refresh.
Security and privacy boundary
Fetch only permitted sources, respect access and content reuse rights, sanitize display content, and protect internal networks from untrusted feed URLs.
Interview follow-ups with reasoning
Question: How do you avoid polling unchanged feeds?
Show answer and explanation
Answer: Use validators such as ETag where supported.
Question: How do you evaluate clustering?
Show answer and explanation
Answer: Review false merges and splits on a labeled sample.
Question: What if ranking is unavailable?
Show answer and explanation
Answer: Serve the last valid snapshot with visible freshness.
Operate and verify the design
Source freshness, malformed-feed rate, false story merges and ranking snapshot age.
Quarantine a bad publisher while proving healthy sources and the last valid feed remain available.
A second scenario to test transfer
A publisher takes 30 seconds to respond while others finish in 200 ms. A single serial fetch loop would let that publisher delay all others. A scheduler with per-host concurrency and a shared bounded worker pool allows unrelated publishers to progress. Retrying the slow host uses backoff, not a tight loop. The user reads an existing feed projection while ingestion continues.
Explain how to recover from a bad ranking deployment without fetching every source again.
Show answer and explanation
Answer: Retain authoritative normalized article facts and revision events. Build a new versioned projection, replay or backfill, catch up changes, validate, and switch the read alias. Keep removal rules effective during migration.
Compare alternatives
| Decision | Benefit | Trade-off |
|---|---|---|
| Precomputed feeds | Fast common reads | Freshness lag and rebuild work |
| Read-time ranking | Personalization flexibility | Higher request cost |
| Per-source scheduling | Failure isolation | More scheduler state |
A design-changing exercise
Two stories describe the same event. Should one be deleted as a duplicate?
Show answer and explanation
Answer: Not necessarily. Distinguish identical syndicated content from independent reporting; preserve provenance and use clustering when appropriate.
Design workshop: normalize, cluster, and rank
Core scope is source ingestion, related-story grouping and a personalized feed. Full-text republication without rights is excluded. Choose source refresh within five minutes for priority sources, feed p95 under 200 ms and corrections/removals visible under an explicit projection bound. Missing a slow source temporarily is different from claiming its article does not exist.
Store sources(id,url,parser_version,etag,last_modified,next_fetch_at,health), articles(id,source_id,canonical_url,revision,published_at,title,body_fingerprint,state), clusters(id,model_version) and user_preferences(user,topics,blocked_sources). Identity first deduplicates canonical URL/source revision; related stories can share a cluster while retaining separate publisher attribution.
A transparent baseline similarity algorithm tokenizes titles/body excerpts, removes boilerplate and compares shingle sets by Jaccard intersection/union. A={storm,city,flood}, B={storm,city,flood,roads} gives 3/4=0.75; C={city,election,vote} gives 1/5=0.2 against A. Require a compatible topic/time window and an illustrative threshold 0.7. Similarity is not proof that two reports are identical; related perspectives remain separate records.
Figure — Clustering without collapsing attribution.
Candidate generation selects recent clusters matching preferences plus a bounded exploration pool. A simple score can be 0.5×topic_match +0.3×freshness +0.2×source_quality, with all inputs normalized and a source-diversity cap. At ages zero and four hours under freshness=exp(-age/4h), freshness is 1 and about 0.368. The weights are original interview assumptions; evaluate click satisfaction and source diversity, not only engagement.
Figure — Rebuild ranking while ingestion continues.
GET /feed returns sessionId, rankVersion, items and signed nextCursor. Freeze a bounded ordered candidate list for five minutes; pagination rechecks removal/blocked-source state. Expiry resets the session. A ranking model change creates a new session rather than reorder page two invisibly. The read path uses precomputed clusters/candidates and bounded personalization; it does not fetch publishers synchronously.
Per-host scheduler uses ETag/Last-Modified when supported and bounded concurrency. A 30-second publisher cannot occupy every worker: cap per-source requests, back off failures and keep independent sources progressing. Store parser version and raw authorized feed evidence for reprocessing. Corrected publication times do not silently overwrite original ingestion evidence.
Exercise: A faulty ranking version groups unrelated election and storm stories. Must the system recrawl all sources to recover?
Show answer and explanation
Answer: No if normalized article facts, revisions and retained processing evidence are authoritative. Build a new clustering/ranking projection, catch up changes and switch the alias after validation. Keep removal rules active during migration.
HTTP conditional requests provides the source-fetch protocol reference; the similarity and ranking policy here is deliberately explicit and illustrative.