Learning pathsA
Patterns

Multi-step Processes

A business process that crosses databases or external services cannot usually be rolled back like one local transaction.

A business process that crosses databases or external services cannot usually be rolled back like one local transaction. Model it as durable state with explicit transitions, retries, and repair. The important question is what happens after a crash between steps, especially when the previous step may already have created an external effect.

Learning goals

Durable state machines; orchestration; outbox; sagas; idempotent activities; compensation; timeouts; reconciliation; manual repair.

The mechanism at a glance

Durable workflow → Inventory hold (step 1); Inventory hold → Payment authorize (step 2); Payment authorize → Confirm order (success); Payment authorize → Release / refund (failure / cancellation); Release / refund → Repair queue (unresolved); Repair queue → Durable workflow (reconcile)
Scroll to inspect the diagram, or open it at full size.

Figure — Durable workflow → Inventory hold (step 1); Inventory hold → Payment authorize (step 2); Payment authorize → Confirm order (success); Payment authorize → Release / refund (failure / cancellation); Release / refund → Repair queue (unresolved); Repair queue → Durable workflow (reconcile)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. Persist the workflow state

Give the workflow an identity, current state, step attempts, and deadline. Store enough information to resume after process loss. A queue message can wake a worker, but the queue alone should not be the only record of the business process. Make transition predicates explicit so duplicate wake-ups do not repeat completed steps blindly.

2. Protect every activity boundary

Reserve inventory using a stable reservation identity, authorize payment using a stable payment identity, and confirm only under valid state. Each activity needs idempotency or reconciliation according to its destination. The workflow engine cannot retroactively make an arbitrary external API exactly once.

3. Use compensation deliberately

If payment authorization fails, release inventory. If a charge completed but fulfillment cannot proceed, a refund may be required. These are new actions with their own failure modes and audit records, not time travel. Some steps are irreversible or only partially compensable; route them to manual repair when needed.

4. Bound time and expose stuck work

Set step deadlines, retry budgets, and escalation policies. Backoff transient failures, quarantine permanent validation problems, and monitor age in each state. A workflow that retries forever without visibility is not reliable from the customer’s perspective. Cancellation must define which already-completed effects remain.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
Order workflow
pending -> inventory_held -> payment_authorized -> confirmed
payment failure -> releasing_inventory -> cancelled
unknown payment -> reconciling -> authorized or failed
Every transition: expected_state + durable attempt identity

Worked example

A worker authorizes payment and crashes before saving the result. On restart it sees the authorization step unfinished. It must query or retry under the same provider identity, not create a new authorization. After resolving the outcome, it records progress and proceeds. This is the central difference between retrying computation and retrying a business effect.

Failure walkthrough

Compensation can fail too. Releasing inventory may time out, or refunding may be temporarily unavailable. Store compensation state and retry it with a stable identity. Show the customer an honest pending resolution status. Reconciliation compares the internal workflow with provider records to find outcomes callbacks never delivered.

Payment authorization sent → Provider accepts → Worker crashes before recording → Resume same workflow step → Resolve using original identity
Scroll to inspect the diagram, or open it at full size.

Figure — Payment authorization sent → Provider accepts → Worker crashes before recording → Resume same workflow step → Resolve using original identity

Decisions and trade-offs

ApproachStrengthCost
OrchestrationVisible central process stateCoordinator and workflow evolution
Event choreographyLooser service couplingHarder end-to-end diagnosis
Saga compensationBusiness recovery across ownersNot equivalent to atomic rollback

Check your understanding

Why is a refund not a rollback of the original charge?

Show answer and explanation

Answer: The charge already existed outside the local transaction and may have observable consequences. A refund is another provider operation that can be delayed, fail, or require reconciliation. Persist and audit both actions.

Transfer to a new scenario

Reserve inventory, authorize a payment, then confirm; recover after each transition.

Why is refunding money a new business action rather than rolling back a completed external call?

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Persist progress across independent systems

A database transaction cannot atomically include an unrelated payment provider and shipping service. Model the business process as durable states with transitions, stable operation IDs and compensations. A compensation is a new business action, such as refunding a charge, and can itself fail or need manual handling.

The transactional outbox solves one specific gap: update local business state and record the intent to publish in the same transaction. A relay publishes later and may publish more than once. Consumers still need deduplication or idempotent transitions. The outbox does not make every downstream side effect exactly once.

Choose orchestration when a central workflow makes dependencies, deadlines and recovery easier to understand. Choreography can decouple producers but makes a long business process harder to trace. In either case, propagate a correlation ID and expose unresolved compensation and uncertain effects to operators.

A decision worksheet for Multi-step Processes: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Multi-step Processes: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
created -> inventory_reserved -> payment_authorized -> confirmed
payment_failed -> release_inventory -> cancelled
unknown_payment -> reconcile -> confirmed or compensate
outbox row commits with each local transition

A tempting mistake

Do not compensate an operation merely because its response timed out. Determine whether it happened, and make compensation safe to repeat. A retry policy without a terminal escalation path can leave workflows stuck forever.

Transfer exercise

What if inventory release succeeds but its acknowledgement is lost?

Show answer and explanation

Answer: Retry the same compensation identity. The release operation checks the reservation state and must not increase available stock twice.

Your study notes