Learning pathsA
GUIDED PRACTICE

Review: Dropbox

Design file upload and synchronization with separate metadata and bytes.

Understand the product before drawing boxes

Imagine saving a video on your laptop, losing the connection halfway through, and reopening the application on another device. The design must distinguish an unfinished transfer from a usable file. Now add a shared folder and two people editing offline. Storage is only one part of the problem: the service must coordinate ownership, completed versions, permissions, and device state without silently losing work.

This is an original interview design, not a reconstruction of Dropbox's production architecture. We use PostgreSQL for authoritative metadata and an S3-compatible object store for bytes. The reasoning matters more than those product names.

Requirements and scope

RequirementObservable behaviorDesign consequence
UploadAn authorized user can transfer a file and resume after interruptionDurable upload session; verified parts; explicit publication
DownloadA user retrieves exactly the authorized versionImmutable object identity; permission check before issuing a grant
ShareA recipient sees shared files and can access allowed operationsIndexed grants with roles and revocation
SyncDevices eventually converge after reconnectingDurable change journal; device cursor; safe local application
Preserve workConcurrent offline edits do not silently overwrite each otherBase-version comparison; conflict copies

For this exercise, support files up to 50 GB and a million uploads per day. Choose a regional metadata API availability target of 99.9% and a p99 metadata-response target of 200 ms under the assumed workload. These are illustrative objectives to discuss, not production promises. A large transfer's duration is dominated by available bandwidth; do not promise the same latency as a metadata lookup. Define a healthy-network sync objective separately, for example notification within five seconds of publication.

Keep browser-based editing, previews, malware scanning, and sophisticated folder inheritance outside the first design. Direct file sharing is included. Historical user-facing version browsing is optional, but immutable internal versions are still useful for safe publication and conflict recovery. Availability does not mean accepting contradictory head updates during a metadata partition: we can delay visibility across regions while serializing updates to one file's head.

Estimate the workload that changes the design

At 1 million uploads/day and 20 MB average size, incoming data is 20 TB/day, approximately 231 MB/s averaged across the day. A fivefold peak is about 1.16 GB/s. Metadata at 2 KB per new version is only 2 GB/day before indexes and replicas. Putting the entire upload through the application servers makes them carry the expensive traffic without improving authorization.

For one 50 GB decimal file on a 100 Mb/s connection, the theoretical transfer time is 50 × 10^9 × 8 / 100 × 10^6 = 4,000 seconds, about 67 minutes. Protocol overhead and interruptions make it longer. Four parallel requests share that connection; they do not create four times its physical capacity.

Choose 16 MiB parts for this example. A 50 GB file needs approximately 2,981 parts. A failure near the end can retry one part instead of the entire file. Part size trades request overhead and bookkeeping against retry cost and memory. Check the selected provider's part-size and part-count limits before making this a production configuration. Do not return thousands of signed URLs in one response; request a bounded batch as the transfer advances.

Entities, identity, and access paths

A file ID identifies a logical record across rename and moves. A version identifies immutable content for that record. An upload ID identifies an attempt to construct a version. A hash identifies bytes; it does not identify their owner or authorize access. Two people can own distinct records with identical content.

sql
files(file_id PK, owner_id, parent_id, name, current_version, deleted_at)
versions(file_id, version, object_key, byte_size, checksum, created_at)
  PRIMARY KEY(file_id, version)
uploads(upload_id PK, file_id, owner_id, base_version, object_key,
        provider_upload_id, state, expires_at, result_version)
upload_parts(upload_id, part_number, etag, checksum, byte_size)
  PRIMARY KEY(upload_id, part_number)
grants(file_id, recipient_id, role, revoked_at)
  PRIMARY KEY(file_id, recipient_id)
changes(namespace_id, sequence, file_id, version, change_type)
  PRIMARY KEY(namespace_id, sequence)

Index files by parent_id plus a stable pagination key for folder listing. Index active grants by recipient_id and file_id for the shared-with-me screen. Listing must not scan every file's recipient array. Version lookup uses the composite primary key. Changes are read by namespace and sequence range. Scope every query to the authenticated user's accessible namespace; a globally unique ID does not bypass permissions.

For the initial regional design, allocate journal positions through a namespace counter locked in the same transaction as the change. This serializes publication in one namespace but gives a clear ordering boundary. A bare sequence allocated before commit is insufficient: transaction 12 can commit before transaction 11, causing a client that advances to 12 to skip 11. At higher load, use a carefully ordered journal or partitioned cursors rather than casually removing the serialization guarantee.

API contracts a client can actually use

The server derives the user from a verified session. An Idempotency-Key is scoped to that user and operation; reuse with a different payload returns a conflict.

json
POST /uploads
Idempotency-Key: laptop-attempt-42
{"fileId":"f17","baseVersion":7,"name":"demo.mp4",
 "byteSize":50000000000,"partSize":16777216}

201
{"uploadId":"u42","state":"pending","nextPartBatch":"/uploads/u42/parts"}

POST /uploads/u42/parts
{"partNumbers":[1,2,3,4]}
200
{"parts":[{"partNumber":1,"url":"short-lived signed PUT URL"}]}

GET /uploads/u42
200
{"state":"transferring","verifiedParts":[1,2,3],"expiresAt":"..."}

POST /uploads/u42/complete
{"parts":[{"partNumber":1,"etag":"..."}],"checksum":"..."}
200
{"fileId":"f17","version":8,"state":"published"}

The completion example abbreviates the full ordered part manifest. Missing parts return a validation error; an unauthorized session returns 403; a changed base version returns 409 with the current version and conflict-resolution options. A published upload returns the same result on retry. An expired session returns an explicit expired response and a restart option. A status endpoint lets a client resolve an ambiguous timeout instead of assuming the operation failed.

Contract / pseudocode
POST /files/f17/download-grants {version:8} -> {url,expiresAt,byteSize,checksum}
POST /files/f17/grants {recipientId,role:"reader"} -> durable grant
DELETE /files/f17/grants/user9 -> revoke future access
GET /shared-files?cursor=... -> authorized records and nextCursor
GET /namespaces/n3/changes?cursor=... -> events,nextCursor,hasMore

Build the architecture in three steps

First consider one application server storing files on its local disk. The request is easy to follow, but another server cannot necessarily read that disk; a failed host can make the file unavailable. Adding more HTTP servers does not repair the storage ownership problem.

Next keep metadata in PostgreSQL and stream bytes through the application to durable object storage. Host failure no longer destroys the only content copy, but each transfer occupies application bandwidth and connections. The database and object store still do not share a normal SQL transaction. A successful object write followed by a failed database write needs recovery.

Finally keep authorization and coordination in the metadata API while the client uploads directly to object storage. The API creates a restricted transfer session and signs permission to upload specific parts at a specific key. Object storage receives the bytes; the API verifies completion and publishes the version. A signed URL is a scoped capability, not a reason to make the entire bucket public.

Three architecture choices and the bottleneck removed by each change.
Scroll to inspect the diagram, or open it at full size.

Upload: from a selected file to a published version

  1. The API checks write permission and records an upload session with an immutable destination key. The previous version remains visible.
  2. The client obtains a small batch of part URLs and sends bytes directly to object storage. Each part has a number and provider response identity.
  3. The client persists its session ID locally. On reconnect, it asks the API which parts the provider actually has and obtains fresh URLs for remaining parts.
  4. The API validates ownership, expected part sizes, part ordering, and integrity. Client progress reports are useful for the interface but are not proof of stored bytes.
  5. The API asks the provider to assemble the object and verifies successful completion. Only then can it publish metadata.
  6. A database transaction compares baseVersion, inserts the immutable version, advances the file head, records the upload result, and appends the ordered change event.
Upload coordination and the direct byte path, with metadata publication after verification.
Scroll to inspect the diagram, or open it at full size.

Do not mark a file usable just because progress reaches 100%. Distinguish bytes transferred, object assembled, and version published in the interface. If the completion response is lost, the client polls status and retries with the same upload identity. The API must return the original version rather than create version 9 for the same completed attempt.

Use the provider's explicit checksum facilities for integrity. An ETag is an opaque protocol value for this contract; do not universally treat it as an MD5 hash, particularly with multipart uploads or encryption. Validate a supported full-object checksum or an explicitly defined part-checksum scheme. Listing parts may require pagination. Provider event notifications are recovery hints, not the sole source of truth for upload completion.

Resume and cleanup without losing good parts

Consider four parts: 1 and 2 are durable, part 3's response is lost, and part 4 has never been sent. The client thinks only two parts succeeded. Provider verification can discover that part 3 exists, so the resumed transfer sends only part 4. If part 3 is absent or fails the expected checksum, retransmit it. A local filename match alone cannot prove that the resumed bytes are the same file; retain content identity and detect modifications during transfer.

Durable part state after a disconnect and the verified resume decision.
Scroll to inspect the diagram, or open it at full size.

Abandoned multipart uploads consume storage. Give sessions a deadline, abort expired provider uploads, and retain enough state to distinguish completed objects from abandoned attempts. Cleanup checks live upload and version references before deleting. A grace period reduces races but is not a substitute for reference checks or exclusion from an active publication.

Download: permission first, bytes through the delivery path

The user asks for file f17 version 8. The API checks the current permission, resolves the immutable object key, and returns a short-lived grant. The client then retrieves bytes from the object store or a private CDN. On a CDN miss, the edge retrieves the immutable object from its protected origin; on a hit, it avoids that origin transfer. The metadata API does not proxy every byte.

Download authorization is separate from the private CDN byte path.
Scroll to inspect the diagram, or open it at full size.

For interrupted downloads, use HTTP byte ranges against the same immutable version and verify the completed content. Upload part boundaries need not be exposed to the downloader. If storage uses independently compressed content blocks, range semantics need an explicit block manifest and reconstruction protocol; a single compressed object does not automatically support seeking in the original byte space.

Caching a private file requires an access check at the delivery boundary. A shared cache must not make one user's authorized response available anonymously. Define revocation honestly: removing a grant stops issuing new links, but an already issued bearer URL can remain usable until expiry unless the delivery layer checks current authorization or supports another revocation mechanism. Deleting downloaded bytes from someone else's device is outside this guarantee.

Sharing: model the queries and revocation

Alice grants Bob read access to f17. Insert a grant keyed by file and recipient. Bob's shared-files screen queries the recipient index, while a download check queries the exact grant and verifies its role. Permission to read does not imply permission to replace content or create more grants. The initial design allows only the owner to share and revoke.

One grant supports both access checks and recipient file listing.
Scroll to inspect the diagram, or open it at full size.

Avoid maintaining an unbounded recipient array and a separate authoritative reverse cache: every update would need both copies to remain consistent. Keep the durable grants authoritative and treat caches as derived. Permission changes also need to notify affected clients. A revoked client may receive a removal event for its local view, but it must not receive later protected metadata or content. Complex shared-folder inheritance needs separate ancestor rules and permission-version invalidation; it is deliberately outside this baseline.

Synchronization in both directions

On the laptop, the sync agent observes local filesystem changes and queues work durably. It coalesces repeated notifications and rechecks file contents; a filesystem event is a hint, and a file can change while it is being read. Upload a stable snapshot or detect the changed local file and retry. Remember the last synchronized version so a local edit has a meaningful baseVersion.

On the phone, a pushed notification says that namespace n3 may have changed. It does not contain the only durable copy of the change. The phone fetches ordered journal events after its cursor, checks permissions, downloads any required version to a temporary file, verifies content, and replaces the local file safely. Advance the cursor only after recording enough durable local state to replay or finish the application.

Publication creates durable changes; notifications wake devices that pull by cursor.
Scroll to inspect the diagram, or open it at full size.

If a notification is lost, periodic polling catches up. If a page of events is fetched twice, apply version-aware idempotent updates. If a cursor expires, return a reset-required response. The client obtains a consistent namespace snapshot paired with a journal position and then consumes changes after that position. An unrelated snapshot followed by a separately sampled position can miss edits in between.

Prevent a feedback loop: applying a remote file creates local filesystem events too. Track the applied remote version and content identity so the agent does not immediately upload the same bytes as a fresh user edit. Rename uses the stable file ID; delete uses a tombstone so disconnected devices learn that absence was intentional.

Two offline edits and a conflict

Laptop and phone both start at version 7. Laptop publishes version 8. Phone later tries to publish its edited contents with baseVersion 7. The server compares the base with the current head inside the publication transaction. It rejects the stale head update and preserves the uploaded object for a bounded resolution period, or creates a separate conflict copy under the stated policy.

A stale base version cannot replace a newer head; preserve both edits.
Scroll to inspect the diagram, or open it at full size.
sql
UPDATE files SET current_version = 8
WHERE file_id = 'f17' AND current_version = 7;
-- Require exactly one updated row, inside the publication transaction.
-- Otherwise roll back publication and return a conflict.

The literal version 8 is illustrative; the application allocates the next valid version. Do not use a client's wall clock to establish ownership of the latest edit. Whole-file conflict copies preserve data but require user resolution. Document collaboration needs a richer merge model and is a different interview problem.

Faster transfer and delta synchronization

Multipart transfer solves resumability; it does not by itself provide persistent cross-version deduplication. A completed multipart object is one object. For delta sync, introduce a content-block layer: immutable chunks addressed by scoped content hashes, an ordered manifest for each file version, and reference management. Upload only missing or changed blocks and commit the manifest after all referenced blocks are durable.

Fixed-size blocks are simple, but inserting bytes near the beginning can shift many later boundaries. Content-defined chunking can reduce that effect by selecting boundaries from content. It adds CPU work, implementation complexity, and variable block sizes. Explain this as an alternative storage model, rather than silently treating multipart part numbers as reusable blocks across every version.

Whole-object transfer and a persistent block manifest are different storage models.
Scroll to inspect the diagram, or open it at full size.

Parallelize only enough to use available bandwidth without excessive memory or retry storms. Compression can help text but often saves little on already compressed media; test total CPU plus transfer time. Hashes do not prove ownership, and cross-tenant deduplication can reveal whether protected content exists. Start with tenant-scoped reuse and authorize every manifest/block request.

Failure windows and recovery

Failure pointDurable stateRecoveryForbidden shortcut
Part response lostProvider may have the partVerify by session and part numberRestart the whole file automatically
Object assembled, API crashesObject exists; file head unchangedReconcile upload status; retry metadata publicationClaim the object and database committed atomically
Metadata committed, response lostVersion and journal event existReturn recorded upload resultPublish a second version for the retry
Notification lostJournal still contains changeDevice polls from saved cursorDepend solely on an open socket
Stale offline editNewer head exists; old-base object may existConflict response/copyOverwrite silently using client time
Cleanup races with completionAn active session may still publishReference checks and coordinated session stateDelete every old unlisted object

Delete first changes authoritative visibility and permissions; physical reclamation follows retention rules. Keep an audit trail of who changed grants and published versions. Replication improves availability but is not a substitute for backups and restore exercises. Test restoring both metadata and referenced content to a consistent recovery point.

Exercise the design and judge your answer

Question: The laptop reports 100% upload progress, but the phone cannot list the new version. What states would you inspect?

Show answer and explanation

Answer: Inspect the provider completion result, upload state, metadata head, and journal position in that order. Transferred parts do not imply a published version. If metadata is committed, inspect the phone's cursor and permissions. Retrying publication is safe only with the original upload identity and recorded result.

Question: Bob loses permission immediately after receiving a five-minute download URL. Can the service promise immediate denial?

Show answer and explanation

Answer: Not with expiration alone. The existing bearer grant may remain usable. Require current authorization at the delivery boundary for a stronger contract, or explicitly advertise a bounded revocation delay.

Question: A one-byte insertion causes nearly every block hash to change. Is the checksum broken?

Show answer and explanation

Answer: Fixed-size boundaries can shift after the insertion. The checksum can be correct while reuse becomes poor. Compare content-defined chunking with its added CPU and manifest costs.

For a foundational answer, demonstrate all four user flows, separate bytes from metadata, and locate publication. A stronger answer explains resume verification, permission queries, cursor ordering, retry identity, and the object/database recovery boundary. For a deeper discussion, quantify tradeoffs in delta sync, namespace ordering, revocation, and storage reclamation. These are study criteria, not universal hiring-level rules.

Measure completion success and time, abandoned bytes, checksum failures, per-device journal lag, permission-revocation delay, and conflict frequency. In a failure drill, disconnect after part 3, lose the completion response, and restart the phone. Success means only missing parts are retransferred, exactly one version is published, and the phone eventually reaches that version without losing its own local edit.

Final architecture and reading connections

The complete design keeps authorization and publication authoritative while bytes bypass the API.
Scroll to inspect the diagram, or open it at full size.

Trace upload, download, share, and sync through this diagram. For each arrow, identify the operation, permission boundary, durable state, and retry identity. A box earns its place only if you can explain its responsibility and its behavior when the adjacent component is unavailable.

Read Amazon S3 multipart upload and checksum semantics for provider-specific mechanisms. Read Dropbox's sync-engine engineering account and compression work for published engineering context. Those accounts describe their own systems; the architecture, numerical assumptions, and failure exercises above are original interview examples.

Design workshop: cleanup, journals, and metadata recovery

The baseline excludes inherited shared-folder permissions. If added, grants must reference ancestors, authorization must evaluate the current ancestor chain, and moving a folder must invalidate permission-version caches. File ownership and an old download capability cannot substitute for those rules.

Use upload_keys(owner_id, request_key, fingerprint, upload_id) with unique owner/request identity. Session creation commits this mapping with upload metadata. A retry returns the same session; changed file size/checksum returns 409. Completion's durable outcome maps the original upload identity to one published version. Expired grants return a refreshable session status rather than forcing retransmission of already verified parts.

A storage calculation separates live objects, retained versions, unfinished multipart data and replicas. At 100,000 users adding 100 MB/day, new logical bytes are 10 TB/day. Ninety days at that rate is 900 TB before version retention, deduplication, compression or provider redundancy. Downloading 200 MB/user/day creates 20 TB/day of delivery traffic. Delta sync saves bandwidth only when the manifest/block model actually supports reused blocks.

Publish versus garbage collection. Publisher to Metadata DB: Claim object generation as live under metadata transaction; GC worker to Metadata DB: Mark unreferenced object generation deleting; Metadata DB to GC worker: Claim succeeds only with no live refs or active upload; GC worker to Object store: Delete that immutable object key; retry safely; Publisher to Metadata DB: Publication cannot attach to a deleting generation; Publisher to Object store: Reupload under a new object generation if needed
Scroll to inspect the diagram, or open it at full size.

Figure — Publish versus garbage collection.

A grace period alone is unsafe: a publisher can gain a reference during the delay. Choose a reference ledger with object states live, gc_candidate and deleting. Publication locks/checks the object row and adds its version reference in the same transaction. GC locks that row and can mark deleting only with zero references and no active publication pins. It never recycles the object key. A crash after physical deletion but before metadata cleanup repeats delete for the same key. Cross-service deletion requires a durable job and reconciliation of metadata versus storage inventories.

Namespace journals serialize change sequence allocation under the same transaction that updates the file head. A counter is therefore a per-namespace contention point. If one namespace sustains 2,000 changes/s and a serial critical section takes 2 ms, a 500/s bound is insufficient even if storage has many shards. Batch independent journal entries under one counter allocation, or explicitly move to a log owner with durable ordering. Sharding files alone does not solve the namespace sequence bottleneck.

Journal recovery and migration checkpoints. Normal reconnect / Cursor still in retained journal / Ordered suffix after saved sequence; Cursor expired / Consistent namespace snapshot + high watermark / Reset then consume events after watermark; Shard migration / Fenced old owner + copied state + caught-up log / New epoch; no two writers assign sequences; Metadata failover / Durable commit boundary + higher owner epoch / Retry identity returns original version
Scroll to inspect the diagram, or open it at full size.

Figure — Journal recovery and migration checkpoints.

A snapshot contains namespace state and journal watermark from one consistent database snapshot. Replay starts strictly after that watermark. During shard migration, copy a snapshot, catch up the suffix, fence the old writer, drain to a final sequence, then activate the new epoch. Requests carrying an old epoch retry through routing; two writers must never independently allocate the same sequence.

For this interview design choose synchronous metadata durability within the primary region, with acknowledged transactions recoverable during a replica failure. An asynchronous remote disaster replica has nonzero RPO; measure its lag and explain the potential lost metadata interval. Physical bytes surviving remotely do not prove their associated committed version survives. If zero-RPO regional failure is required, metadata commits need the stronger multi-region quorum contract and its latency/availability cost. Set recovery objectives explicitly, for example a five-minute regional recovery target subject to tested failover, not a promise derived from the diagram.

Exercise: Snapshot watermark is 500, a file change commits as 501 during snapshot transfer, and a wake-up notification is lost. Is the change lost?

Show answer and explanation

Answer: No if the snapshot and watermark were consistent and the client pulls the durable journal after 500. Notification is only a hint. If sequence 501 has already aged out, return a reset requirement and another consistent snapshot; do not silently jump to the newest cursor.

Amazon S3 multipart upload documents the provider boundary. Reference management, version publication and journal ordering remain application responsibilities.

26:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

When can an uploaded version appear in another device’s change stream?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.