Learning pathsA
GUIDED PRACTICE

Handling Large Blobs

Large files should not usually pass through application workers that only need to authorize and track them.

Large files should not usually pass through application workers that only need to authorize and track them. Separate the control path—identity, permissions, metadata, and state—from the byte path between clients, object storage, and delivery edges. This keeps API capacity tied to decisions rather than the duration of multi-gigabyte transfers.

Learning goals

Separate metadata and bytes; signed URLs; multipart; content checks; finalization; stale uploads; access control; CDN; lifecycle cleanup.

The mechanism at a glance

Client → Control API (permission); Client → Object storage (large byte path); Object storage → Verifier (completed object); Verifier → Metadata ready (publish after checks); Metadata ready → CDN / download (authorized delivery); CDN / download → Client (bytes)
Scroll to inspect the diagram, or open it at full size.

Figure — Client → Control API (permission); Client → Object storage (large byte path); Object storage → Verifier (completed object); Verifier → Metadata ready (publish after checks); Metadata ready → CDN / download (authorized delivery); CDN / download → Client (bytes)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. Authorize a bounded transfer

Create an upload session for a specific user, object identity, size policy, and expiry. Issue narrowly scoped signed credentials. A signed URL is a bearer capability while valid, so avoid logging or exposing it casually. The application must still verify the final object rather than assuming the client obeyed metadata constraints.

2. Transfer and verify

Multipart upload permits parallel parts and targeted retries. Track part identities and integrity information, then complete the upload according to the storage API. Choose checksums deliberately; an object ETag is not universally the checksum your application expects. Bound client and service concurrency to avoid overwhelming the network or storage API.

3. Publish a complete version

Keep pending and complete states distinct. Finalization checks ownership, expected bytes, integrity, and any required scanning before making the object visible. Use immutable object keys per version so a stale upload cannot overwrite the current file. A database transaction cannot atomically roll back an already completed object-store transfer.

4. Deliver and clean up

Authorize downloads and issue appropriately scoped access or route through an edge policy. Define revocation behavior for already-issued URLs and cached content. Abort stale multipart sessions and garbage-collect unreferenced objects with a grace period. Lifecycle cleanup is part of the design, not an optional future task.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
POST /uploads -> {upload_id, part_authorizations}
Client -> object storage: bytes
POST /uploads/:id/finalize -> verify -> complete metadata
GET /files/:id/download -> authorization -> limited access
States: pending, validating, complete, rejected, expired

Worked example

A user uploads a 5 GB recording over an unreliable connection. The application creates one session and the client retries only failed parts. Application workers do not hold the entire file in memory. After storage completion, the verifier checks the expected object and marks it ready. If processing is required, ready-for-processing and ready-for-playback remain separate states.

Failure walkthrough

The upload credential expires while some parts are unfinished. Authenticate a request to renew the permitted session rather than mint unrestricted access. A cleanup task must distinguish abandoned sessions from active resumable work. If download access is revoked, explain the remaining validity of previously issued links; short expiry bounds exposure but is not instant recall.

Session created pending → Some parts uploaded → Network interruption → Resume missing parts → Verify then publish complete
Scroll to inspect the diagram, or open it at full size.

Figure — Session created pending → Some parts uploaded → Network interruption → Resume missing parts → Verify then publish complete

Decisions and trade-offs

BoundaryMechanismWhy it matters
Control / bytesDirect scoped transferAvoid long-lived API workers
Pending / completeVerified finalizationNo partial file visibility
Version identityImmutable object keyNo stale overwrite
CleanupGrace period + referencesNo permanent orphan growth

Check your understanding

What prevents a partially uploaded file from appearing as complete?

Show answer and explanation

Answer: An explicit metadata state transition after object completion and verification. Readers only expose complete versions. The presence of some uploaded parts or a client success message is insufficient.

Transfer to a new scenario

Upload a 5 GB recording without holding an application worker for the whole transfer.

What prevents an unfinished multipart upload from becoming a visible downloadable file?

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Separate byte transfer from metadata ownership

Application servers should usually authorize a transfer and issue a scoped upload session while object storage accepts the bytes. Track expected size, checksum, part identities and expiry. The completion operation validates the assembled object before making a metadata version visible. A signed URL is a capability whose scope and lifetime must be limited.

Multipart upload makes retry cost proportional to failed parts. Chunk size trades number of requests against wasted retransmission and parallelism. Content-addressed chunks can deduplicate data, but cross-user deduplication may reveal content existence and complicate deletion. Immutable version keys simplify caches and retries.

Garbage collection must distinguish an orphan from an upload still in progress. Use session expiry, a grace period and reference checks. For downloads, decide whether byte-range requests, CDN caching and resumable clients are needed. Revocation guarantees must match token and cache lifetime.

A decision worksheet for Handling Large Blobs: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Handling Large Blobs: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
create upload session -> scoped part URLs
upload parts -> record checksums
complete -> verify manifest -> publish metadata version
expire abandoned session -> grace period -> collect objects

A tempting mistake

Completing an object-store upload and committing application metadata are separate boundaries. A crash can leave an orphan; trying to make the object visible first can expose incomplete or unauthorized content.

Transfer exercise

A client retries complete after losing the response. How do you avoid two versions?

Show answer and explanation

Answer: Bind completion to the upload ID and store its committed version. A retry returns that version after rechecking authorization.

7:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

What prevents a partially uploaded file from appearing as complete?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.