Learning pathsA
Key Technologies

Temporal

Temporal is a workflow platform for durable application processes that must wait, retry, and resume across worker failure.

Temporal is a workflow platform for durable application processes that must wait, retry, and resume across worker failure. Use it when process state and time span matter more than a single request handler. It does not make external side effects exactly once; activities and their destinations still need idempotency and reconciliation.

The mechanism at a glance

Workflow start → Durable event history (stable workflow ID); Durable event history → Task queue (schedule activity); Task queue → Worker activity (poll); Worker activity → External service (idempotent effect); External service → Durable event history (record outcome); Signal / timer → Durable event history (resume state)
Scroll to inspect the diagram, or open it at full size.

Figure — Workflow start → Durable event history (stable workflow ID); Durable event history → Task queue (schedule activity); Task queue → Worker activity (poll); Worker activity → External service (idempotent effect); External service → Durable event history (record outcome); Signal / timer → Durable event history (resume state)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. 1 · Frame

Model the business workflow first: identity, state transitions, signals, timers, cancellation, and terminal conditions. A workflow execution replays its event history, so workflow code must obey deterministic execution rules. Put network calls, database access, and other side effects in activities. Keep workflow histories bounded through appropriate continuation strategies for very long or high-event processes.

2. 2 · Model

Start workflows with stable IDs so duplicate start requests resolve intentionally. Activities have task queues, timeouts, and retry policies. Choose activity idempotency keys derived from workflow and step identity. Signals deliver external events into workflow state; validate their identity and expected sequence. Child workflows can isolate repeated sub-processes, but their state and limits remain operational concerns.

3. 3 · Scale

Workers poll task queues, so scale them according to activity duration and external capacity. Temporal service stores execution histories; operators need visibility into backlog, retries, task schedules, and stuck workflows. Version workflow behavior as deployed code evolves because replay may encounter histories created by older logic. Define cancellation and compensation as product behavior, not just a worker shutdown.

4. 4 · Recover

A workflow activity charges a provider and times out after the charge succeeds but before the result is recorded. Retry with the same provider identity or query status. Workflow history records orchestration progress, not every fact inside an external system. Reconciliation remains necessary for missing callbacks, provider state drift, and manual repair.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
Workflow(id, type, input_version, state history)
Activity(step_id, attempt, timeout, retry policy)
ExternalEffect key = workflow_id + step_id
Signal(event_id, expected_transition)
History: deterministic decisions; side effects happen in activities

Worked example

An order workflow reserves inventory, waits up to 15 minutes for payment, then confirms or releases. The process can remain durable while no worker is active during the wait. When payment signal arrives, the workflow validates its identity and advances. If the worker restarts, it replays recorded decisions and schedules only the activity still required.

Failure walkthrough

A deployment changes workflow code so replaying a prior history follows a different command sequence. That can fail determinism checks. Introduce compatible workflow versioning and migration strategy, then test old histories before rolling out. Do not force-reset live executions without understanding business state or external effects already completed.

Workflow waits for payment → Timer is persisted → Worker deployment restarts → History replays deterministically → Payment signal resumes next step
Scroll to inspect the diagram, or open it at full size.

Figure — Workflow waits for payment → Timer is persisted → Worker deployment restarts → History replays deterministically → Payment signal resumes next step

Decisions and trade-offs

Use Temporal whenSimpler alternativeTrade-off
Multi-step durable processOne queue + job rowWorkflow history and versioning
Long timers / signalsScheduled task tableService dependency and operational model
Activity retriesApplication retry loopDuplicate external effect risk remains
Human waitPolling a status fieldLong-lived execution and retention

Check your understanding

An activity times out after an external charge might have completed. What should the retry do?

Show answer and explanation

Answer: Use an idempotent operation identity stable across activity attempts, or query the provider by the original reference before issuing a new charge. The workflow can durably retry orchestration, but it cannot infer the external result from its timeout alone.

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Continue the connection

Study Managing Long Running Tasks and explain which guarantee from this lesson carries into that topic.

Distinguish replayed workflow code from retried activities

A durable workflow records its history and reconstructs logical state by replay. Workflow code must obey the platform’s deterministic execution rules; arbitrary network calls, unrecorded randomness or uncontrolled wall-clock reads can make replay diverge. Activities perform external work and have their own retry and timeout behavior.

If a payment activity times out, the workflow must not assume the provider did nothing. Give the activity a stable business operation key and reconcile uncertain outcomes. Activity retries can rerun code after partial completion. Durable orchestration removes much scheduling and state bookkeeping, but it does not remove the need for idempotent side effects.

Long histories need a lifecycle strategy such as continuing into a new run while preserving necessary state. Code changes must remain compatible with in-flight histories using supported versioning techniques. Separate activity queue delay, execution time, heartbeat progress and total workflow deadline when configuring limits.

A decision worksheet for Temporal: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Temporal: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
workflow: reserve -> authorize -> confirm
activity input includes stable payment_operation_id
on uncertainty: query provider before new charge
on business failure: schedule retryable compensation
record workflow version and operational deadline

A tempting mistake

A workflow retry that creates a new business key can charge again. The identity of the business operation should not be regenerated merely because a worker restarted.

Transfer exercise

Why must workflow code avoid directly reading an external API during replay?

Show answer and explanation

Answer: The response can change and replay must reconstruct recorded decisions consistently. External effects belong in activities or another supported recorded mechanism.

Your study notes