Skip to content
Navigation
Dashboard
← All system design problems
Easy · Social Media · About 45 minutes

Design Instagram - Photo Sharing Platform

Understand the requirements. Trace the requests. Explain the trade-offs.

CANDIDATE-LED INTERVIEW WALKTHROUGH

Design Instagram: publish safely, build the feed and handle fan-out

A candidate-led photo-publishing and following-feed design, with an explicit evolution to ranking and celebrity traffic. The proposed design is not a claim about Instagram internals.

Interviewer replies, workloads and targets are illustrative assumptions to agree on in an interview. Use this as a study resource: establish scope, draw a complete baseline, then choose the most consequential deep dives with your interviewer.

1. Ask which Instagram we are designing

Candidate explains

Instagram includes several products. I would ask whether this is photo publishing and the following feed, or Reels, Stories, Explore, messaging or live video. I will begin with publishing and a following feed.

Feed order, audience privacy and follower skew are the most consequential questions; they matter more than naming a cache early.

Candidate asksIllustrative interviewer replyDesign consequence
Which user journeys?Publish photos, follow accounts, read feed, like and comment.Five complete paths; defer DMs, ads and live video.
Chronological or ranked feed?Start chronological; explain ranking later.Correct eligible candidates before ML complexity.
Private accounts and blocking?Yes, including revocation.Recheck current visibility at read time.
Follower distribution?Most users have hundreds; some creators have millions.No unbounded synchronous celebrity fan-out.
Freshness target?New posts appear within seconds for most followers.Asynchronous fan-out with measurable lag/replay.
Media scope?Photos with bounded sizes first.Validate and create renditions; defer full video pipeline.

2. Define scope and publication correctness

Candidate explains

I will support durable publication, follow/unfollow, paginated reads, likes and comments. Reverse chronological order is the baseline. Ranking later reorders eligible candidates; it must not change authorization rules.

A post becomes visible only when required approved renditions are ready and metadata commits. Feed entries are references, not an authority for content visibility. A stale feed cache cannot overrule deletion or privacy.

RequirementTarget or rule
Feed latencyIllustrative first-page p95 below 300 ms, excluding media bytes.
Freshnessp95 ordinary-author fan-out lag below 5 seconds at admitted load; measure celebrity cases separately.
Availability99.9% monthly feed target; degrade ranking/counts before authorization.
PrivacyCheck current audience/block/deletion state before returning content access; fail closed when required authority is unavailable.
ConsistencyPublish post and outbox atomically; feed placement and counts may lag.
ExclusionsReels, Stories expiry, ads auction, DMs and live video are separate extensions.

3. Show how follower skew changes capacity

Candidate explains

Assume 10M DAU, 20 feed pages per user/day and 20 posts/page. Assume 1M new photo posts/day and 200 eligible followers per ordinary author. I explicitly separate the celebrity tail; averages hide its operational impact.

CalculationResultDesign consequence
10M × 20 pages ÷ 86,4002,315 reads/s average; 23,148/s at assumed 10× peakCache candidate IDs and batch hydration.
1M posts × 200 followers/day200M candidate inserts/day; 2,315/s averageAsynchronous ordinary-author fan-out is plausible.
One post × 10M followers10M potential inserts at oncePull popular-author candidates or selectively fan out in bounded jobs.
1M × (2 MB original + 1 MB renditions)3 TB/day; 90 TB per 30 days before replicasSeparate media retention from feed indexes.
200M pages × 20 thumbnails × 100 KB400 TB/day illustrative image transferCDN and lazy loading dominate bandwidth; not every image needs a full-resolution fetch.

4. Separate source data from the derived feed

Candidate explains

I begin with indexed relational metadata for users, follows, posts, likes and comments, plus private object storage. The feed index contains bounded post references and can be rebuilt.

A like is a unique relationship, not an unprotected increment. Retried operations must not inflate counts; displayed counters can be derived and reconciled.

EntityAccess patternInvariant
Post(postId,authorId,time,audience,mediaVersion,state)Author/time and IDOnly published, ready media enters the feed.
Follow(followerId,followeeId,approvalState)Both graph directionsPrivate approval and blocks evaluated explicitly.
FeedEntry(viewerId,postId,publishTime)Viewer reverse-time rangeUnique viewer/post pair; bounded and rebuildable.
Like(userId,postId), Comment(commentId,postId,author,time)Unique like; comments by post/timeIdempotent mutations with authorization.
ProcessingJob(uploadId,ownerId,objectKey,pipelineVersion,state)Status and retry workSame output version on retry; poison input quarantined.

5. Define the client contract

Candidate explains

Clients upload through scoped capabilities, but a client saying “complete” does not publish a post. A server pipeline validates content and produces required renditions.

Pagination uses an opaque stable ordering cursor. Offsets drift when new posts arrive. A ranked feed needs an explicit ranking-session/snapshot policy rather than re-sorting every page independently.

InterfaceRequest → resultRetry / error contract
POST /v1/media/uploadsType, size, checksum → uploadId and scoped URLOwner-scoped key and enforced limits.
POST /v1/postsuploadId, caption, audience → postId and PROCESSING/READYRepeat same key returns same post; never publish incomplete output.
PUT /v1/follows/{accountId}Desired follow state → active or approval-pendingIdempotent with block/privacy checks.
GET /v1/feed?cursor=…&limit=20Eligible hydrated posts and next cursorBounded work, current authorization and deduplication.
PUT /v1/posts/{id}/likes/meDesired like state → current stateUnique relationship; retry is not another increment.
POST /v1/posts/{id}/commentsclientCommentId and text → commentIdDeduplicate and authorize against current post state.

6. Start with a complete baseline

Candidate explains

A simple correct feed pulls recent posts from followed authors and merges by publication time. Publishing is cheap, but read work grows with follow count. After measuring that bottleneck I add a derived inbox for ordinary authors.

Both versions share the same authoritative published-post/visibility rules. I do not begin with a large ranking platform before establishing reliable publication and candidate delivery.

7. Trace upload, processing and publication

Candidate explains

Transfer completion, processing completion and visible publication are distinct states. I make each retryable and protect publication against stale jobs. A delete racing a worker must never make content reappear.

  1. Authorize transfer

    Check quota and create an owner-specific random object key with a short-lived capability. Do not allow overwrite of published content.

  2. Validate/process

    Verify size/type/checksum, isolate decoding, strip sensitive metadata according to policy and create required renditions. Untrusted input cannot run arbitrary processing work.

  3. Publish atomically

    After outputs pass checks, transition PROCESSING to PUBLISHED and append an outbox event in the same transaction. Check pipeline/post version so an old job cannot regress state.

  4. Fan out in pages

    Enumerate eligible follower partitions in bounded jobs. Checkpoint and use unique viewer/post insertions so retries do not duplicate or skip items.

  5. Serve safely

    Hydrate current state, check visibility and issue appropriately bounded content access. Filter stale candidates.

  6. Delete and reconcile

    Tombstone source data first. Asynchronously clean feed references/media according to retention. Track stuck processing and fan-out jobs.

8. Merge feed candidates and handle celebrity posts

Candidate explains

For a hybrid feed I load ordinary-author inbox references, fetch bounded recent windows for followed popular authors, merge/deduplicate, check visibility and hydrate a page. Ranking is optional and follows candidate eligibility.

The celebrity threshold is a measured cost decision based on active followers, posting/read rates and lag. One universal follower count is not a defensible threshold.

  1. Bound candidates

    Cap inbox and per-author windows. Isolate slow sources with budgets; prevent a user following many accounts from triggering unbounded work.

  2. Merge and paginate

    Deduplicate by postId and use (publishTime,postId) in a stable read window. New posts appear on refresh without repeating old page items.

  3. Authorize

    Evaluate current private-account approval, blocking, post audience and tombstones. Cleanup is not the privacy enforcement mechanism.

  4. Hydrate

    Batch author/post/count reads and lazy-load media. Previously issued media grants have a declared lifetime; downloaded plaintext cannot be recalled.

  5. Evolve ranking

    Retrieve features, score and diversify eligible candidates with final safety checks. On ranker timeout return a chronological eligible fallback.

9. Explain alternatives and degradation

Candidate explains

The feed derivation is disposable and observable. I can degrade ranking or counters, but I cannot bypass a required authorization check to fill the page.

AlternativeBenefitCost
Fan-out on readCheap publication and natural celebrity handlingRead amplification for users following many authors.
Fan-out on writeFast bounded readsWrite amplification, offline-user waste and lag.
HybridPrecompute ordinary posts and merge popular authorsMore complex pagination, deduplication and observability.
Cached countersFast responseEventual counts need unlike/delete reconciliation.
RankingPotentially better relevanceFeature/model latency, feedback loops and pagination stability require measurement.

10. Test user-visible failures

Candidate explains

I measure publish success/age, fan-out lag by author class, first-page latency, media failures, duplicate entries, privacy denials and ranking fallback rate. Aggregate averages can hide a broken celebrity workflow.

FailureRecoveryAssertion
Worker crashes after one renditionRetry same pipeline versionNo incomplete published media.
Fan-out stops midwayResume checkpoint; unique insertsNo duplicate or permanently skipped recipient partition.
Celebrity publishes at peakPull or bounded selective jobsOrdinary publication does not starve.
Private access is revokedAuthoritative read/capability checkStale inbox cannot grant new access.
Ranker times outChronological eligible fallbackUsable feed without bypassing privacy.
Delete races with late workerVersion/state-conditional publishDeleted content never reappears.

11. Summarize the design and choose the next deep dive

Candidate explains

I covered publish, follow, feed, likes and comments. Publication is authoritative; the derived feed can lag or be rebuilt. Hybrid fan-out addresses skew that average QPS hides.

Senior depth includes retries, privacy, deletion and end-to-end read/write flow. Staff depth adds graph ownership, hot-author isolation, regional behavior, rebuild strategy and measured ranking quality within latency/safety budgets.

Technical references

Primary references explain underlying mechanisms. Workloads and architecture choices above remain proposed interview assumptions.