Learning pathsA
GUIDED PRACTICE

News Aggregator

Design a personalized list of articles collected from publishers.

Interview scope and guarantees

Ingest feeds, normalize stories, group duplicates, rank recent items and return a paginated feed. Start with public syndicated sources. Define freshness, attribution and what happens when publishers correct or retract a story.

Capacity worksheet

Assume 100,000 sources polled every 15 minutes: about 111 polls/s before retries. One million articles/day at 10 KB is 10 GB/day raw. Ten million users reading 5 times/day generate 579 reads/s average; cache popular category feeds rather than reranking the full corpus for every request.

Concrete API contract

Contract / pseudocode
POST /sources {feedUrl,category}
GET /feed?category=...&cursor=...
GET /stories/{id}
POST /stories/{id}/feedback {action}

Data model and access paths

Contract / pseudocode
sources(id PK,url,etag,next_poll_at,last_success)
articles(id PK,source_id,canonical_url,published_at,content_hash,version)
clusters(cluster_id,article_id,similarity_version)
feed_generation(category,generation,ordered_story_ids)

Evolve a solution and explain each change

Three architecture decisions for News Aggregator, including the pressure each introduces.
Scroll to inspect the diagram, or open it at full size.

Figure — Three architecture decisions for News Aggregator, including the pressure each introduces.

Step 1: Collect source articles

Fetch feeds under source limits and preserve provenance. The same story can arrive with different URLs and titles.

Step 2: Normalize and index

Assign article identity, cluster related stories, and build read projections. Similarity is not proof that two reports are interchangeable.

Step 3: Serve personalized pages

Rank bounded candidates with freshness and source diversity, using stable cursors. Re-ranking a changing list can duplicate pages or hide updates.

Responsibility overview

Connected responsibilities for News Aggregator. Trace the authoritative and derived paths separately.
Scroll to inspect the diagram, or open it at full size.

Figure — Connected responsibilities for News Aggregator. Trace the authoritative and derived paths separately.

Worked end-to-end scenario

Two feeds publish the same syndicated article under different URLs. URL normalization catches only some duplicates; content and publisher metadata can help identify a common article. A different reporter’s story about the same event belongs in the same cluster but should keep its own identity and attribution. Fetchers persist source versions before indexing. A corrected headline updates the article version, and the projection rejects an older replay. A reader page uses a ranking snapshot or cursor contract so an inserted breaking story does not shift offsets unpredictably. If the projection is unavailable, a simpler recent-articles view can be a defined degradation.

Why these access paths matter

Store source URL, publisher, fetched version, timestamps, and licensing/provenance attributes. Index articles by publication time and category; cluster membership is a revisable projection. User preferences belong in a separate access-controlled record. Keep fetching cadence per source and distinguish ingestion lag from ranking freshness.

Build the baseline first

News Aggregator: baseline request paths.
Scroll to inspect the diagram, or open it at full size.
  1. Polite poller → raw feed → normalized articles.
  2. Deduplication → story clusters → ranker.
  3. Published feed generation → cache → reader.

Evolve the design under load

News Aggregator: additional scaling and recovery paths.
Scroll to inspect the diagram, or open it at full size.
  1. Source queues → retry isolation → bounded fetchers.
  2. Incremental ranking → category snapshots.
  3. Feedback stream → bounded personalization features.

Defend the hardest decision

Canonical URLs and exact hashes catch obvious duplicates; semantic clustering can group related reporting but may incorrectly merge different events. Retain original article identities and attribution even when presenting a story cluster. Publication time supplied by a feed may be wrong or updated, so use ingestion time for pipeline monitoring and a documented ranking policy for freshness.

Failure and recovery analysis

One publisher begins returning malformed or enormous XML. Apply parser limits, timeouts and quarantine to that source without blocking ingestion globally. If an article is retracted, propagate a versioned tombstone through clusters and cached feeds. A ranked snapshot gives stable pagination while newer generations remain available on refresh.

Security and privacy boundary

Fetch only permitted sources, respect access and content reuse rights, sanitize display content, and protect internal networks from untrusted feed URLs.

Interview follow-ups with reasoning

Question: How do you avoid polling unchanged feeds?

Show answer and explanation

Answer: Use validators such as ETag where supported.

Question: How do you evaluate clustering?

Show answer and explanation

Answer: Review false merges and splits on a labeled sample.

Question: What if ranking is unavailable?

Show answer and explanation

Answer: Serve the last valid snapshot with visible freshness.

Operate and verify the design

Source freshness, malformed-feed rate, false story merges and ranking snapshot age.

Quarantine a bad publisher while proving healthy sources and the last valid feed remain available.

A second scenario to test transfer

A publisher takes 30 seconds to respond while others finish in 200 ms. A single serial fetch loop would let that publisher delay all others. A scheduler with per-host concurrency and a shared bounded worker pool allows unrelated publishers to progress. Retrying the slow host uses backoff, not a tight loop. The user reads an existing feed projection while ingestion continues.

Explain how to recover from a bad ranking deployment without fetching every source again.

Show answer and explanation

Answer: Retain authoritative normalized article facts and revision events. Build a new versioned projection, replay or backfill, catch up changes, validate, and switch the read alias. Keep removal rules effective during migration.

Compare alternatives

DecisionBenefitTrade-off
Precomputed feedsFast common readsFreshness lag and rebuild work
Read-time rankingPersonalization flexibilityHigher request cost
Per-source schedulingFailure isolationMore scheduler state

A design-changing exercise

Two stories describe the same event. Should one be deleted as a duplicate?

Show answer and explanation

Answer: Not necessarily. Distinguish identical syndicated content from independent reporting; preserve provenance and use clustering when appropriate.

Design workshop: normalize, cluster, and rank

Core scope is source ingestion, related-story grouping and a personalized feed. Full-text republication without rights is excluded. Choose source refresh within five minutes for priority sources, feed p95 under 200 ms and corrections/removals visible under an explicit projection bound. Missing a slow source temporarily is different from claiming its article does not exist.

Store sources(id,url,parser_version,etag,last_modified,next_fetch_at,health), articles(id,source_id,canonical_url,revision,published_at,title,body_fingerprint,state), clusters(id,model_version) and user_preferences(user,topics,blocked_sources). Identity first deduplicates canonical URL/source revision; related stories can share a cluster while retaining separate publisher attribution.

A transparent baseline similarity algorithm tokenizes titles/body excerpts, removes boilerplate and compares shingle sets by Jaccard intersection/union. A={storm,city,flood}, B={storm,city,flood,roads} gives 3/4=0.75; C={city,election,vote} gives 1/5=0.2 against A. Require a compatible topic/time window and an illustrative threshold 0.7. Similarity is not proof that two reports are identical; related perspectives remain separate records.

Clustering without collapsing attribution. A and B / 3 shared of 4 union = 0.75 / Same related-story cluster under chosen threshold; A and C / 1 shared of 5 union = 0.20 / Different cluster; Identical syndicated body / Matching normalized fingerprint / Duplicate presentation candidate; retain provenance; Correction with same URL / New source revision / Update article facts and versioned cluster/rank projection
Scroll to inspect the diagram, or open it at full size.

Figure — Clustering without collapsing attribution.

Candidate generation selects recent clusters matching preferences plus a bounded exploration pool. A simple score can be 0.5×topic_match +0.3×freshness +0.2×source_quality, with all inputs normalized and a source-diversity cap. At ages zero and four hours under freshness=exp(-age/4h), freshness is 1 and about 0.368. The weights are original interview assumptions; evaluate click satisfaction and source diversity, not only engagement.

Rebuild ranking while ingestion continues. Source fetcher to Article DB: Commit normalized article revision and source provenance; New projector to Article DB: Snapshot facts at watermark W; New projector to New projector: Recluster/rank under version 4; Source fetcher to Article DB: New edits/removals append after W; New projector to Article DB: Catch up suffix; apply tombstones before switch; Feed API to New projector: Read alias switches to verified projection version 4
Scroll to inspect the diagram, or open it at full size.

Figure — Rebuild ranking while ingestion continues.

GET /feed returns sessionId, rankVersion, items and signed nextCursor. Freeze a bounded ordered candidate list for five minutes; pagination rechecks removal/blocked-source state. Expiry resets the session. A ranking model change creates a new session rather than reorder page two invisibly. The read path uses precomputed clusters/candidates and bounded personalization; it does not fetch publishers synchronously.

Per-host scheduler uses ETag/Last-Modified when supported and bounded concurrency. A 30-second publisher cannot occupy every worker: cap per-source requests, back off failures and keep independent sources progressing. Store parser version and raw authorized feed evidence for reprocessing. Corrected publication times do not silently overwrite original ingestion evidence.

Exercise: A faulty ranking version groups unrelated election and storm stories. Must the system recrawl all sources to recover?

Show answer and explanation

Answer: No if normalized article facts, revisions and retained processing evidence are authoritative. Build a new clustering/ranking projection, catch up changes and switch the alias after validation. Keep removal rules active during migration.

HTTP conditional requests provides the source-fetch protocol reference; the similarity and ranking policy here is deliberately explicit and illustrative.

Technical references

Kafka processing and external-sink boundaries.

PostgreSQL transaction isolation and concurrent updates.

10:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

Explain how to recover from a bad ranking deployment without fetching every source again.

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.