Design Dropbox: from clarification to recoverable file synchronization
Follow a candidate through direct transfers, verified publication, sharing, offline edits and recovery. The architecture and workloads are proposed interview assumptions, not company implementation claims.
1. Ask before drawing
Before selecting storage, I would ask whether you want the file-storage product or the internals of an object store. I will design the product unless you prefer the latter. I will ask the questions that change the design, state assumptions where you want me to proceed, and avoid trying to recreate every feature.
| Candidate asks | Illustrative interviewer reply | Design consequence |
|---|---|---|
| Which actions are core? | Upload, download, share and sync across devices. | Four complete user journeys; editing and search are excluded. |
| How large are files and how reliable are clients? | Up to 50 GB; mobile connections can drop. | Direct resumable transfer and durable upload-session progress. |
| Can two devices edit offline? | Yes; do not silently lose either edit. | Compare base versions and preserve conflict copies. |
| What does sharing mean? | Named users, view/edit permissions and revocation. | Indexed ACLs and an explicit capability-expiry window. |
| What scale should I use? | State a reasonable illustrative workload. | One internally consistent ledger, not unrelated company statistics. |
| Can remote devices briefly lag? | Yes, but publication and authorization must be correct. | Strong metadata mutations; eventual notification delivery. |
2. Agree on scope, targets and invariants
I will support upload/resume, download, named-user sharing and synchronization. I will retain versions for conflict recovery, but defer collaborative editing, previews, global search, public links and cross-user deduplication.
The central invariant is that every published version points to verified, readable content. A successful object upload is not yet a published file. An object ID is never proof of permission.
| Requirement | Agreed interpretation |
|---|---|
| Metadata latency | Illustrative p95 below 200 ms in-region at admitted load; bulk transfer time is separate. |
| Sync freshness | p95 notification delay below 5 seconds for connected devices; excludes downloading a large file. |
| Availability | 99.9% monthly metadata API target; reject unsafe version writes during an ownership failure. |
| Durability | Replicate metadata before acknowledging, retain immutable objects and test restores; storage-provider durability is not an end-to-end guarantee. |
| Consistency | Transactional version/ACL mutations; other devices and notifications may lag. |
| Security | TLS, encrypted storage, per-operation ACL checks, bounded signed capabilities, quotas and abuse controls. |
3. Work from a single ledger
I use decimal GB/TB/PB unless a binary unit is explicitly written. Assume 10M daily active users, 100M current files averaging 10 MB, and 2M new versions/day averaging 5 MB. These are design inputs. The different average sizes describe stored files versus daily changes.
The byte path dominates transfer capacity, while metadata handles small records. That distinction justifies direct storage transfers before it justifies adding many microservices.
| Calculation | Result | Design consequence |
|---|---|---|
| 100M files × 10 MB | 1 PB current logical content | Versions, replication and backups are additional. |
| 2M versions ÷ 86,400 | 23.1 uploads/s average; 231/s at assumed 10× peak | Size session/commit traffic separately from byte traffic. |
| 2M versions × 5 MB | 10 TB/day; 115.7 MB/s average | Approximately 1.16 GB/s at the assumed peak, before overhead. |
| 10M DAU × 100 metadata calls/day | 11,574 calls/s average | Measure owner skew and read/write mix before partitioning. |
| 50 GB × 8 ÷ 100 Mb/s | 4,000 seconds ideal transfer time | Resume is essential; no API latency target can remove transfer time. |
| 50 GB ÷ 8 MiB logical chunks | About 5,961 chunks | Bound/page manifest and session responses. |
4. Model files, versions and ownership
I start with a relational metadata store for atomic version publication, folder changes and ACL updates. Immutable objects hold content. A key-value design could also work, but would require an equally explicit conditional-write and transaction story.
A file, version, manifest and upload session are different entities. In this design logical chunks are separate immutable objects, allowing reuse across versions. They are not S3 multipart parts, which become a single object after completion.
| Entity | Access pattern | Correctness rule |
|---|---|---|
| File(fileId, ownerId, parentId, name, currentVersion, deletedAt) | List by owner/parent/name; lookup by stable ID | Unique live name per parent; rename preserves identity. |
| FileVersion(fileId, versionId, baseVersion, manifestId) | Ordered versions per file | Immutable; currentVersion changed with compare-and-swap. |
| Manifest(ordered chunk IDs, sizes, checksums) | Fetch by manifestId | Represents exact byte order, not a set of hashes. |
| UploadSession(ownerId, sessionId, baseVersion, state, expiresAt) | Authorized session status; expiry index | Owner-scoped idempotency, quota reservation and content verification. |
| Share(fileId, granteeId, role) | Index both file and grantee | Current ACL is authoritative. |
| Change(ownerId, cursor, event) and outbox | Sequential replay per owner | Append atomically with metadata publication. |
5. Define APIs, identity and retries
Identity comes from authentication, not a trusted userId in the body. Each operation checks current authorization. I bind an idempotency key to its caller and request digest; reusing the key with different input returns a conflict.
Clients need to distinguish expired sessions, missing content and stale versions. A timeout is an unknown outcome, not permission to blindly create another version.
| Interface | Request → response | Recovery contract |
|---|---|---|
| POST /v1/uploads | Destination, baseVersion, size, manifest digest → sessionId and paginated scoped chunk URLs | Same operation key replays the recorded session; enforce quota and size limits. |
| GET /v1/uploads/{id} | Authorized lookup → verified chunk status and expiry | Refresh URLs only for the same authorized upload. |
| POST /v1/uploads/{id}/commit | Manifest and baseVersion → fileId/versionId | 409 stale base; incomplete content unpublished; repeat returns stored result. |
| GET /v1/files/{id}/download | Optional version → manifest and short-lived URLs | Recheck access; renewal requires authorization. |
| GET /v1/changes?cursor=…&limit=100 | Owner-scoped cursor → ordered events and next cursor | Expired cursor returns resync-required, never silent success. |
| PUT /v1/files/{id}/shares/{grantee} | Role or revocation → ACL version | Only an authorized owner/admin can mutate permissions. |
6. Draw a small complete design
I can begin with a metadata application, replicated database, object storage and a worker. The control plane authorizes operations; the client moves bulk bytes directly. A durable change log/outbox connects committed metadata to retryable notifications.
A CDN may cache immutable content, but the origin is private and content capabilities expire. I do not need a separate service for every noun to explain a correct baseline.
7. Trace upload and the durable acknowledgement
Let me trace a laptop uploading an edit. The success boundary is the metadata transaction publishing the verified manifest and change event. It is not the final object PUT.
If objects succeed and metadata fails, I temporarily have unreachable objects to collect. That is preferable to publishing a version with missing bytes. The object store and database do not share an atomic transaction.
Prepare locally
Read a stable file snapshot, compute ordered chunks/checksums and maintain a local journal. If the file changes while reading, create a fresh version attempt rather than mixing two edits.
Create the session
Authorize the destination, reserve quota and record baseVersion plus the operation key. Only reuse content the owner is already authorized to reference; a global hash-existence API can leak other users’ data.
Transfer and resume
Upload missing chunks with bounded parallelism and backoff. After a crash query session state, refresh expired URLs and continue missing chunks. Reuse the session identity.
Verify
Check object identity, size and supported integrity metadata server-side. A client assertion is not evidence. An ETag is not universally a content hash.
Publish
Atomically compare currentVersion to baseVersion, insert the immutable version/manifest, advance the visible pointer, append a change and save the commit result. A lost response is recovered by replay.
Notify and clean up
Retry the outbox and collect only proven unreferenced content after a safety interval. Cleanup must coordinate with publication/reference ownership; it cannot race-delete a committed chunk.
8. Explain download, sync, conflicts and sharing
A push notification only says that something changed. Devices pull the durable change log with a stored cursor after reconnect and periodically as a safety net. A socket cannot guarantee no missed changes.
For downloads, authorize a specific version, verify chunks, assemble a temporary local file and atomically replace the destination. Advance the cursor only after applying or durably journaling the batch.
Offline conflict
Two devices edit version 7. The first publishes version 8. The second base-version comparison fails; preserve its edit as a conflict copy rather than overwriting based on device clocks.
Rename/delete
Use stable IDs and tombstones. Rename cannot detach in-flight content. Keep tombstones for the supported offline window; older cursors require full reconciliation.
Sharing
List shared files through a grantee index rather than scanning all files. Commit ACL changes and their change/outbox records together.
Revocation
Deny new capabilities after the authoritative ACL update. Previously issued URLs may work until expiry, and already downloaded plaintext cannot be recalled. State that product boundary clearly.
9. Evolve the baseline deliberately
I partition metadata by owner or shared workspace and keep version changes within the owning partition. I measure hot folders separately from averages. Cross-owner sharing needs explicit authorization relationships; it is not a free distributed transaction.
Fixed-size chunks are simple, but an insertion near the beginning can shift many subsequent boundaries. Content-defined chunking is a deliberate extension with bounded sizes and a versioned algorithm, not something automatically supplied by multipart upload.
| Pressure | Evolution | Trade-off |
|---|---|---|
| Hot downloads | Cache immutable versions near clients | Egress/cache efficiency versus capability-expiry and privacy policy. |
| Repeated edits | Content-defined chunking and owner-scoped reuse | Extra client CPU and manifest bookkeeping. |
| Metadata load | Indexed listings, owner partitions and caches | Never make stale ACL cache data the authority. |
| Regional outage | Replicated objects/metadata and exercised recovery | Declare RPO/RTO and any unsafe-write rejection period. |
| Slow client | Adaptive bounded parallel chunk transfers | Cannot exceed available network bandwidth; cap memory use. |
10. Test partial failure, not only the happy path
I inject faults between transfer, verification, metadata commit and notification. The assertion is one recoverable visible result with no missing or unauthorized bytes. Track oldest upload age, publish failure rate, sync lag, conflict rate, orphan backlog and restore time.
| Failure | Recovery | Assertion |
|---|---|---|
| Client dies after most uploads | Resume the owner-scoped session | Completed chunks reused; expired quota released safely. |
| Commit succeeds; reply lost | Replay stored result | One version and one logical change, not two. |
| Notification worker crashes | Retry outbox and replay cursor log | No permanent missing change; duplicates harmless. |
| Cleanup races with commit | Ownership/reachability checks plus grace interval | No referenced chunk deleted. |
| Object replication lags at failover | Honor declared recovery point and gate unsafe reads/writes | Never declare complete recovery while referenced content is absent. |
| Access revoked after URL issue | Deny renewal and bound old capability lifetime | Explicit exposure window; no promise to erase downloaded plaintext. |
11. Close with the guarantees
My design uses immutable content, transactional visible versions and a durable change log. Upload, download, sharing and synchronization each have a complete flow. I can go deeper into conflict resolution, object lifecycle or regional recovery.
A senior answer should explain commit/replay boundaries, authorization and offline conflicts. A staff discussion adds ownership migration, hot shared workspaces, measurable recovery objectives and operational validation. More boxes do not replace those guarantees.
Technical references
Primary references explain underlying mechanisms. Workloads and architecture choices above remain proposed interview assumptions.