Managing Long Running Tasks
Long tasks outlive ordinary HTTP requests and often outlive the process that started them.
Long tasks outlive ordinary HTTP requests and often outlive the process that started them. The API should promise durable acceptance and provide a way to inspect progress. Reliability comes from explicit state, recoverable ownership, checkpoints, and bounded retries—not from holding a connection open until something eventually finishes.
Learning goals
202 acceptance; job IDs; durable state; leases; heartbeats; checkpoints; retries; cancellation; deadlines; result ownership; fairness.
The mechanism at a glance
Figure — 202 acceptance → Durable job (commit intent); Durable job → Leased worker (attempt token); Leased worker → Checkpoint store (progress); Leased worker → Fenced completion (conditional result); Cancel request → Durable job (state race); Durable job → Fenced completion (current state)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Create a durable job contract
POST /jobs validates the request, stores a job and outbox event, and returns 202 with a status URL. Record the input version, owner, deadline, and idempotency identity. The status endpoint distinguishes queued, running, succeeded, failed, and cancelled. Progress is informative but should not be mistaken for a committed result.
2. Lease and checkpoint execution
A worker claims an attempt token and renews a lease while active. Checkpoint expensive progress under that attempt’s ownership. If the worker disappears, another can resume from a verified checkpoint or restart. Choose checkpoint frequency by balancing repeated work against checkpoint cost.
3. Fence completion and cancellation
A lease can expire while a paused worker is still alive. Publish results only with the current token and allowed state transition. Cancellation also needs an explicit race rule: if cancellation commits first, later success publication is rejected; if success commits first, cancellation reports that the task already completed. External effects may require compensation.
4. Manage fairness and deadlines
Separate short interactive jobs from large batch tasks, apply tenant quotas, and bound retries. Retry transient infrastructure failures, but do not repeatedly run a permanently invalid input. A deadline should stop useless work and trigger a visible terminal or repair state. Measure oldest job age and time spent in each stage.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
Job(id, owner, input_version, state, deadline, current_token)
Attempt(job_id, token, lease_until, checkpoint_ref)
POST /jobs/:id/cancel -> conditional transition
Publish result WHERE state='running' AND token=current_tokenWorked example
An export worker finishes generating a file just as the user cancels. If cancellation wins the state transition, the worker must not expose a success link afterward. Its uploaded object becomes cleanup work. If success wins first, the cancellation endpoint returns the completed status. A single authoritative state transition gives the interface a coherent answer even when the two events are nearly simultaneous.
Failure walkthrough
A worker heartbeats successfully but makes no useful progress because one dependency hangs. Heartbeats show liveness, not progress. Track stage duration and checkpoint advancement, apply dependency deadlines, and enforce the job’s overall deadline. When retries exhaust the budget, retain diagnostic state and allow an intentional user or operator retry with clear semantics.
Figure — Worker uploads output → User requests cancellation → Cancel transition commits → Worker tries success transition → Reject result; schedule cleanup
Decisions and trade-offs
| Mechanism | Purpose | Limitation |
|---|---|---|
| Lease | Recover abandoned ownership | Old worker may still run |
| Checkpoint | Reduce repeated work | Must match input and attempt |
| Fencing | Protect final state | Destination must enforce it |
| Deadline | Bound useless work | External effects need cleanup |
Check your understanding
An old worker uploads a result after cancellation. What should become visible to the user?
Show answer and explanation
Answer: The authoritative cancelled state should remain if it won the transition. Reject stale result publication and clean up the unreferenced object. Uploading bytes alone must not change the job’s terminal state.
Transfer to a new scenario
Cancel an export while a worker is uploading output and define which terminal result wins.
What stops a worker that lost its lease from committing a stale result?
Continue the connection
Study Multi-step Processes and explain which guarantee from this lesson carries into that topic.
Track ownership and checkpoints independently
Accept a task as a durable resource and return its ID. Workers claim a lease, checkpoint progress and publish terminal state. The queue is a delivery mechanism; the task record explains what the user requested and which attempt currently owns completion. Keep attempt history so retries do not erase evidence.
A checkpoint should represent completed, reusable work. For a video job, that may be a verified chunk output; for a bulk export, a source snapshot and last emitted key. Saving an arbitrary in-memory offset without the associated output commit can skip or duplicate data after recovery.
Cancellation is a request, not a guarantee that every effect vanished. Workers check it at safe boundaries, stop new work and clean up temporary output. Hard termination limits resource use, while conditional result commits prevent a killed or expired attempt from overwriting a newer outcome.
Figure — A decision worksheet for Managing Long Running Tasks: read the mechanism and its guarantee together.
Operational sketch
job(id,state,generation,lease_until,checkpoint)
claim increments generation
heartbeat extends only current generation
complete WHERE generation=:mine AND state="running"
cancel records intent; worker acknowledges cancellationA tempting mistake
A heartbeat can continue while useful work is stuck. Monitor progress age separately from lease age, and bound retries by elapsed time, attempts and business deadline.
Transfer exercise
What if the worker uploads its result then loses the lease before committing?
Show answer and explanation
Answer: The conditional commit fails. Keep the result attempt-specific and let garbage collection remove it unless a current attempt explicitly adopts and verifies it.
An old worker uploads a result after cancellation. What should become visible to the user?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.