Handling Large Blobs
Large files should not usually pass through application workers that only need to authorize and track them.
Large files should not usually pass through application workers that only need to authorize and track them. Separate the control path—identity, permissions, metadata, and state—from the byte path between clients, object storage, and delivery edges. This keeps API capacity tied to decisions rather than the duration of multi-gigabyte transfers.
Learning goals
Separate metadata and bytes; signed URLs; multipart; content checks; finalization; stale uploads; access control; CDN; lifecycle cleanup.
The mechanism at a glance
Figure — Client → Control API (permission); Client → Object storage (large byte path); Object storage → Verifier (completed object); Verifier → Metadata ready (publish after checks); Metadata ready → CDN / download (authorized delivery); CDN / download → Client (bytes)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Authorize a bounded transfer
Create an upload session for a specific user, object identity, size policy, and expiry. Issue narrowly scoped signed credentials. A signed URL is a bearer capability while valid, so avoid logging or exposing it casually. The application must still verify the final object rather than assuming the client obeyed metadata constraints.
2. Transfer and verify
Multipart upload permits parallel parts and targeted retries. Track part identities and integrity information, then complete the upload according to the storage API. Choose checksums deliberately; an object ETag is not universally the checksum your application expects. Bound client and service concurrency to avoid overwhelming the network or storage API.
3. Publish a complete version
Keep pending and complete states distinct. Finalization checks ownership, expected bytes, integrity, and any required scanning before making the object visible. Use immutable object keys per version so a stale upload cannot overwrite the current file. A database transaction cannot atomically roll back an already completed object-store transfer.
4. Deliver and clean up
Authorize downloads and issue appropriately scoped access or route through an edge policy. Define revocation behavior for already-issued URLs and cached content. Abort stale multipart sessions and garbage-collect unreferenced objects with a grace period. Lifecycle cleanup is part of the design, not an optional future task.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
POST /uploads -> {upload_id, part_authorizations}
Client -> object storage: bytes
POST /uploads/:id/finalize -> verify -> complete metadata
GET /files/:id/download -> authorization -> limited access
States: pending, validating, complete, rejected, expiredWorked example
A user uploads a 5 GB recording over an unreliable connection. The application creates one session and the client retries only failed parts. Application workers do not hold the entire file in memory. After storage completion, the verifier checks the expected object and marks it ready. If processing is required, ready-for-processing and ready-for-playback remain separate states.
Failure walkthrough
The upload credential expires while some parts are unfinished. Authenticate a request to renew the permitted session rather than mint unrestricted access. A cleanup task must distinguish abandoned sessions from active resumable work. If download access is revoked, explain the remaining validity of previously issued links; short expiry bounds exposure but is not instant recall.
Figure — Session created pending → Some parts uploaded → Network interruption → Resume missing parts → Verify then publish complete
Decisions and trade-offs
| Boundary | Mechanism | Why it matters |
|---|---|---|
| Control / bytes | Direct scoped transfer | Avoid long-lived API workers |
| Pending / complete | Verified finalization | No partial file visibility |
| Version identity | Immutable object key | No stale overwrite |
| Cleanup | Grace period + references | No permanent orphan growth |
Check your understanding
What prevents a partially uploaded file from appearing as complete?
Show answer and explanation
Answer: An explicit metadata state transition after object completion and verification. Readers only expose complete versions. The presence of some uploaded parts or a client success message is insufficient.
Transfer to a new scenario
Upload a 5 GB recording without holding an application worker for the whole transfer.
What prevents an unfinished multipart upload from becoming a visible downloadable file?
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Separate byte transfer from metadata ownership
Application servers should usually authorize a transfer and issue a scoped upload session while object storage accepts the bytes. Track expected size, checksum, part identities and expiry. The completion operation validates the assembled object before making a metadata version visible. A signed URL is a capability whose scope and lifetime must be limited.
Multipart upload makes retry cost proportional to failed parts. Chunk size trades number of requests against wasted retransmission and parallelism. Content-addressed chunks can deduplicate data, but cross-user deduplication may reveal content existence and complicate deletion. Immutable version keys simplify caches and retries.
Garbage collection must distinguish an orphan from an upload still in progress. Use session expiry, a grace period and reference checks. For downloads, decide whether byte-range requests, CDN caching and resumable clients are needed. Revocation guarantees must match token and cache lifetime.
Figure — A decision worksheet for Handling Large Blobs: read the mechanism and its guarantee together.
Operational sketch
create upload session -> scoped part URLs
upload parts -> record checksums
complete -> verify manifest -> publish metadata version
expire abandoned session -> grace period -> collect objectsA tempting mistake
Completing an object-store upload and committing application metadata are separate boundaries. A crash can leave an orphan; trying to make the object visible first can expose incomplete or unauthorized content.
Transfer exercise
A client retries complete after losing the response. How do you avoid two versions?
Show answer and explanation
Answer: Bind completion to the upload ID and store its committed version. A retry returns that version after rechecking authorization.