Design Instagram: publish safely, build the feed and handle fan-out
A candidate-led photo-publishing and following-feed design, with an explicit evolution to ranking and celebrity traffic. The proposed design is not a claim about Instagram internals.
1. Ask which Instagram we are designing
Instagram includes several products. I would ask whether this is photo publishing and the following feed, or Reels, Stories, Explore, messaging or live video. I will begin with publishing and a following feed.
Feed order, audience privacy and follower skew are the most consequential questions; they matter more than naming a cache early.
| Candidate asks | Illustrative interviewer reply | Design consequence |
|---|---|---|
| Which user journeys? | Publish photos, follow accounts, read feed, like and comment. | Five complete paths; defer DMs, ads and live video. |
| Chronological or ranked feed? | Start chronological; explain ranking later. | Correct eligible candidates before ML complexity. |
| Private accounts and blocking? | Yes, including revocation. | Recheck current visibility at read time. |
| Follower distribution? | Most users have hundreds; some creators have millions. | No unbounded synchronous celebrity fan-out. |
| Freshness target? | New posts appear within seconds for most followers. | Asynchronous fan-out with measurable lag/replay. |
| Media scope? | Photos with bounded sizes first. | Validate and create renditions; defer full video pipeline. |
2. Define scope and publication correctness
I will support durable publication, follow/unfollow, paginated reads, likes and comments. Reverse chronological order is the baseline. Ranking later reorders eligible candidates; it must not change authorization rules.
A post becomes visible only when required approved renditions are ready and metadata commits. Feed entries are references, not an authority for content visibility. A stale feed cache cannot overrule deletion or privacy.
| Requirement | Target or rule |
|---|---|
| Feed latency | Illustrative first-page p95 below 300 ms, excluding media bytes. |
| Freshness | p95 ordinary-author fan-out lag below 5 seconds at admitted load; measure celebrity cases separately. |
| Availability | 99.9% monthly feed target; degrade ranking/counts before authorization. |
| Privacy | Check current audience/block/deletion state before returning content access; fail closed when required authority is unavailable. |
| Consistency | Publish post and outbox atomically; feed placement and counts may lag. |
| Exclusions | Reels, Stories expiry, ads auction, DMs and live video are separate extensions. |
3. Show how follower skew changes capacity
Assume 10M DAU, 20 feed pages per user/day and 20 posts/page. Assume 1M new photo posts/day and 200 eligible followers per ordinary author. I explicitly separate the celebrity tail; averages hide its operational impact.
| Calculation | Result | Design consequence |
|---|---|---|
| 10M × 20 pages ÷ 86,400 | 2,315 reads/s average; 23,148/s at assumed 10× peak | Cache candidate IDs and batch hydration. |
| 1M posts × 200 followers/day | 200M candidate inserts/day; 2,315/s average | Asynchronous ordinary-author fan-out is plausible. |
| One post × 10M followers | 10M potential inserts at once | Pull popular-author candidates or selectively fan out in bounded jobs. |
| 1M × (2 MB original + 1 MB renditions) | 3 TB/day; 90 TB per 30 days before replicas | Separate media retention from feed indexes. |
| 200M pages × 20 thumbnails × 100 KB | 400 TB/day illustrative image transfer | CDN and lazy loading dominate bandwidth; not every image needs a full-resolution fetch. |
4. Separate source data from the derived feed
I begin with indexed relational metadata for users, follows, posts, likes and comments, plus private object storage. The feed index contains bounded post references and can be rebuilt.
A like is a unique relationship, not an unprotected increment. Retried operations must not inflate counts; displayed counters can be derived and reconciled.
| Entity | Access pattern | Invariant |
|---|---|---|
| Post(postId,authorId,time,audience,mediaVersion,state) | Author/time and ID | Only published, ready media enters the feed. |
| Follow(followerId,followeeId,approvalState) | Both graph directions | Private approval and blocks evaluated explicitly. |
| FeedEntry(viewerId,postId,publishTime) | Viewer reverse-time range | Unique viewer/post pair; bounded and rebuildable. |
| Like(userId,postId), Comment(commentId,postId,author,time) | Unique like; comments by post/time | Idempotent mutations with authorization. |
| ProcessingJob(uploadId,ownerId,objectKey,pipelineVersion,state) | Status and retry work | Same output version on retry; poison input quarantined. |
5. Define the client contract
Clients upload through scoped capabilities, but a client saying “complete” does not publish a post. A server pipeline validates content and produces required renditions.
Pagination uses an opaque stable ordering cursor. Offsets drift when new posts arrive. A ranked feed needs an explicit ranking-session/snapshot policy rather than re-sorting every page independently.
| Interface | Request → result | Retry / error contract |
|---|---|---|
| POST /v1/media/uploads | Type, size, checksum → uploadId and scoped URL | Owner-scoped key and enforced limits. |
| POST /v1/posts | uploadId, caption, audience → postId and PROCESSING/READY | Repeat same key returns same post; never publish incomplete output. |
| PUT /v1/follows/{accountId} | Desired follow state → active or approval-pending | Idempotent with block/privacy checks. |
| GET /v1/feed?cursor=…&limit=20 | Eligible hydrated posts and next cursor | Bounded work, current authorization and deduplication. |
| PUT /v1/posts/{id}/likes/me | Desired like state → current state | Unique relationship; retry is not another increment. |
| POST /v1/posts/{id}/comments | clientCommentId and text → commentId | Deduplicate and authorize against current post state. |
6. Start with a complete baseline
A simple correct feed pulls recent posts from followed authors and merges by publication time. Publishing is cheap, but read work grows with follow count. After measuring that bottleneck I add a derived inbox for ordinary authors.
Both versions share the same authoritative published-post/visibility rules. I do not begin with a large ranking platform before establishing reliable publication and candidate delivery.
7. Trace upload, processing and publication
Transfer completion, processing completion and visible publication are distinct states. I make each retryable and protect publication against stale jobs. A delete racing a worker must never make content reappear.
Authorize transfer
Check quota and create an owner-specific random object key with a short-lived capability. Do not allow overwrite of published content.
Validate/process
Verify size/type/checksum, isolate decoding, strip sensitive metadata according to policy and create required renditions. Untrusted input cannot run arbitrary processing work.
Publish atomically
After outputs pass checks, transition PROCESSING to PUBLISHED and append an outbox event in the same transaction. Check pipeline/post version so an old job cannot regress state.
Fan out in pages
Enumerate eligible follower partitions in bounded jobs. Checkpoint and use unique viewer/post insertions so retries do not duplicate or skip items.
Serve safely
Hydrate current state, check visibility and issue appropriately bounded content access. Filter stale candidates.
Delete and reconcile
Tombstone source data first. Asynchronously clean feed references/media according to retention. Track stuck processing and fan-out jobs.
8. Merge feed candidates and handle celebrity posts
For a hybrid feed I load ordinary-author inbox references, fetch bounded recent windows for followed popular authors, merge/deduplicate, check visibility and hydrate a page. Ranking is optional and follows candidate eligibility.
The celebrity threshold is a measured cost decision based on active followers, posting/read rates and lag. One universal follower count is not a defensible threshold.
Bound candidates
Cap inbox and per-author windows. Isolate slow sources with budgets; prevent a user following many accounts from triggering unbounded work.
Merge and paginate
Deduplicate by postId and use (publishTime,postId) in a stable read window. New posts appear on refresh without repeating old page items.
Authorize
Evaluate current private-account approval, blocking, post audience and tombstones. Cleanup is not the privacy enforcement mechanism.
Hydrate
Batch author/post/count reads and lazy-load media. Previously issued media grants have a declared lifetime; downloaded plaintext cannot be recalled.
Evolve ranking
Retrieve features, score and diversify eligible candidates with final safety checks. On ranker timeout return a chronological eligible fallback.
9. Explain alternatives and degradation
The feed derivation is disposable and observable. I can degrade ranking or counters, but I cannot bypass a required authorization check to fill the page.
| Alternative | Benefit | Cost |
|---|---|---|
| Fan-out on read | Cheap publication and natural celebrity handling | Read amplification for users following many authors. |
| Fan-out on write | Fast bounded reads | Write amplification, offline-user waste and lag. |
| Hybrid | Precompute ordinary posts and merge popular authors | More complex pagination, deduplication and observability. |
| Cached counters | Fast response | Eventual counts need unlike/delete reconciliation. |
| Ranking | Potentially better relevance | Feature/model latency, feedback loops and pagination stability require measurement. |
10. Test user-visible failures
I measure publish success/age, fan-out lag by author class, first-page latency, media failures, duplicate entries, privacy denials and ranking fallback rate. Aggregate averages can hide a broken celebrity workflow.
| Failure | Recovery | Assertion |
|---|---|---|
| Worker crashes after one rendition | Retry same pipeline version | No incomplete published media. |
| Fan-out stops midway | Resume checkpoint; unique inserts | No duplicate or permanently skipped recipient partition. |
| Celebrity publishes at peak | Pull or bounded selective jobs | Ordinary publication does not starve. |
| Private access is revoked | Authoritative read/capability check | Stale inbox cannot grant new access. |
| Ranker times out | Chronological eligible fallback | Usable feed without bypassing privacy. |
| Delete races with late worker | Version/state-conditional publish | Deleted content never reappears. |
11. Summarize the design and choose the next deep dive
I covered publish, follow, feed, likes and comments. Publication is authoritative; the derived feed can lag or be rebuilt. Hybrid fan-out addresses skew that average QPS hides.
Senior depth includes retries, privacy, deletion and end-to-end read/write flow. Staff depth adds graph ownership, hot-author isolation, regional behavior, rebuild strategy and measured ranking quality within latency/safety budgets.
Technical references
Primary references explain underlying mechanisms. Workloads and architecture choices above remain proposed interview assumptions.