Learning pathsA
GUIDED PRACTICE

Managing Long Running Tasks

Long tasks outlive ordinary HTTP requests and often outlive the process that started them.

Long tasks outlive ordinary HTTP requests and often outlive the process that started them. The API should promise durable acceptance and provide a way to inspect progress. Reliability comes from explicit state, recoverable ownership, checkpoints, and bounded retries—not from holding a connection open until something eventually finishes.

Learning goals

202 acceptance; job IDs; durable state; leases; heartbeats; checkpoints; retries; cancellation; deadlines; result ownership; fairness.

The mechanism at a glance

202 acceptance → Durable job (commit intent); Durable job → Leased worker (attempt token); Leased worker → Checkpoint store (progress); Leased worker → Fenced completion (conditional result); Cancel request → Durable job (state race); Durable job → Fenced completion (current state)
Scroll to inspect the diagram, or open it at full size.

Figure — 202 acceptance → Durable job (commit intent); Durable job → Leased worker (attempt token); Leased worker → Checkpoint store (progress); Leased worker → Fenced completion (conditional result); Cancel request → Durable job (state race); Durable job → Fenced completion (current state)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. Create a durable job contract

POST /jobs validates the request, stores a job and outbox event, and returns 202 with a status URL. Record the input version, owner, deadline, and idempotency identity. The status endpoint distinguishes queued, running, succeeded, failed, and cancelled. Progress is informative but should not be mistaken for a committed result.

2. Lease and checkpoint execution

A worker claims an attempt token and renews a lease while active. Checkpoint expensive progress under that attempt’s ownership. If the worker disappears, another can resume from a verified checkpoint or restart. Choose checkpoint frequency by balancing repeated work against checkpoint cost.

3. Fence completion and cancellation

A lease can expire while a paused worker is still alive. Publish results only with the current token and allowed state transition. Cancellation also needs an explicit race rule: if cancellation commits first, later success publication is rejected; if success commits first, cancellation reports that the task already completed. External effects may require compensation.

4. Manage fairness and deadlines

Separate short interactive jobs from large batch tasks, apply tenant quotas, and bound retries. Retry transient infrastructure failures, but do not repeatedly run a permanently invalid input. A deadline should stop useless work and trigger a visible terminal or repair state. Measure oldest job age and time spent in each stage.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
Job(id, owner, input_version, state, deadline, current_token)
Attempt(job_id, token, lease_until, checkpoint_ref)
POST /jobs/:id/cancel -> conditional transition
Publish result WHERE state='running' AND token=current_token

Worked example

An export worker finishes generating a file just as the user cancels. If cancellation wins the state transition, the worker must not expose a success link afterward. Its uploaded object becomes cleanup work. If success wins first, the cancellation endpoint returns the completed status. A single authoritative state transition gives the interface a coherent answer even when the two events are nearly simultaneous.

Failure walkthrough

A worker heartbeats successfully but makes no useful progress because one dependency hangs. Heartbeats show liveness, not progress. Track stage duration and checkpoint advancement, apply dependency deadlines, and enforce the job’s overall deadline. When retries exhaust the budget, retain diagnostic state and allow an intentional user or operator retry with clear semantics.

Worker uploads output → User requests cancellation → Cancel transition commits → Worker tries success transition → Reject result; schedule cleanup
Scroll to inspect the diagram, or open it at full size.

Figure — Worker uploads output → User requests cancellation → Cancel transition commits → Worker tries success transition → Reject result; schedule cleanup

Decisions and trade-offs

MechanismPurposeLimitation
LeaseRecover abandoned ownershipOld worker may still run
CheckpointReduce repeated workMust match input and attempt
FencingProtect final stateDestination must enforce it
DeadlineBound useless workExternal effects need cleanup

Check your understanding

An old worker uploads a result after cancellation. What should become visible to the user?

Show answer and explanation

Answer: The authoritative cancelled state should remain if it won the transition. Reject stale result publication and clean up the unreferenced object. Uploading bytes alone must not change the job’s terminal state.

Transfer to a new scenario

Cancel an export while a worker is uploading output and define which terminal result wins.

What stops a worker that lost its lease from committing a stale result?

Continue the connection

Study Multi-step Processes and explain which guarantee from this lesson carries into that topic.

Track ownership and checkpoints independently

Accept a task as a durable resource and return its ID. Workers claim a lease, checkpoint progress and publish terminal state. The queue is a delivery mechanism; the task record explains what the user requested and which attempt currently owns completion. Keep attempt history so retries do not erase evidence.

A checkpoint should represent completed, reusable work. For a video job, that may be a verified chunk output; for a bulk export, a source snapshot and last emitted key. Saving an arbitrary in-memory offset without the associated output commit can skip or duplicate data after recovery.

Cancellation is a request, not a guarantee that every effect vanished. Workers check it at safe boundaries, stop new work and clean up temporary output. Hard termination limits resource use, while conditional result commits prevent a killed or expired attempt from overwriting a newer outcome.

A decision worksheet for Managing Long Running Tasks: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Managing Long Running Tasks: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
job(id,state,generation,lease_until,checkpoint)
claim increments generation
heartbeat extends only current generation
complete WHERE generation=:mine AND state="running"
cancel records intent; worker acknowledges cancellation

A tempting mistake

A heartbeat can continue while useful work is stuck. Monitor progress age separately from lease age, and bound retries by elapsed time, attempts and business deadline.

Transfer exercise

What if the worker uploads its result then loses the lease before committing?

Show answer and explanation

Answer: The conditional commit fails. Keep the result attempt-specific and let garbage collection remove it unless a current attempt explicitly adopts and verifies it.

7:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

An old worker uploads a result after cancellation. What should become visible to the user?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.