Learning pathsA
Question Breakdowns

Job Scheduler

Design a service that runs delayed and recurring jobs with retries, cancellation, and multiple worker regions.

Interview scope and guarantees

Create one-time or recurring jobs, pause schedules, trigger due occurrences, inspect outcomes and retry failed work. Define timezone and daylight-saving semantics. A schedule is not an execution: each occurrence has its own stable identity.

Capacity worksheet

Assume 10 million schedules with 1 million due per hour: 278/s average. If 100,000 are due at midnight, a one-minute dispatch target requires 1,667/s before execution cost. Partition by due-time bucket and tenant so one synchronized tenant cannot monopolize the fleet.

Concrete API contract

Contract / pseudocode
POST /schedules {expression,timeZone,payload,target}
PATCH /schedules/{id} {expectedVersion,enabled,expression}
GET /runs/{id}
POST /runs/{id}/retry

Data model and access paths

Contract / pseudocode
schedules(id PK,tenant_id,expression,time_zone,version,enabled,next_run_at,misfire_policy)
occurrences(schedule_id,schedule_version,scheduled_time,run_id); UNIQUE(schedule_id,schedule_version,scheduled_time)
runs(id PK,state,generation,next_eligible_at,lease_until); INDEX(state,next_eligible_at)
attempts(run_id,generation,worker_id,started_at,finished_at,outcome); PK(run_id,generation)
outbox(event_id PK,run_id,published_at)

Evolve a solution and explain each change

Three architecture decisions for Job Scheduler, including the pressure each introduces.
Scroll to inspect the diagram, or open it at full size.

Figure — Three architecture decisions for Job Scheduler, including the pressure each introduces.

Step 1: Persist due work

Store job definitions and next eligible execution time. A timer in one process disappears on restart.

Step 2: Lease executions

Claim due runs atomically with attempts, deadlines, and retry policy. Expired workers can finish after a replacement starts.

Step 3: Fence effects and calendars

Separate scheduled runs from attempts; define timezone, missed-run, and overlap rules. A successful task side effect can precede an ambiguous acknowledgement.

Responsibility overview

Connected responsibilities for Job Scheduler. Trace the authoritative and derived paths separately.
Scroll to inspect the diagram, or open it at full size.

Figure — Connected responsibilities for Job Scheduler. Trace the authoritative and derived paths separately.

Worked end-to-end scenario

A daily job has run identity (job J, scheduled time T). Worker W1 leases attempt 1 and calls an external service. Its acknowledgement is lost and the lease expires. W2 retries attempt 2 using the same run-level idempotency identity for the external effect. W1’s later status update is fenced by its old attempt token. Scheduling tomorrow is distinct from retrying today’s run. For a daylight-saving transition, the schedule must say whether a repeated local time yields one run or two and how a nonexistent local time is handled. On a prolonged outage, choose skip, coalesce, or bounded catch-up explicitly.

Why these access paths matter

Index runnable executions by next eligible time and state. A unique job-plus-scheduled-time key prevents duplicate logical runs. Store attempts separately from run outcome and maintain a current fencing token. Partition due work into bounded claim batches; a lease is ownership of an attempt, not proof that the external task has never executed.

Build the baseline first

Job Scheduler: baseline request paths.
Scroll to inspect the diagram, or open it at full size.
  1. Schedule API → durable rule → due-time index.
  2. Due scanner → unique occurrence and outbox.
  3. Queue → leased worker → durable outcome.

Evolve the design under load

Job Scheduler: additional scaling and recovery paths.
Scroll to inspect the diagram, or open it at full size.
  1. Time buckets → sharded scanners → fair queues.
  2. Worker pools → per-target concurrency budgets.
  3. Overdue scanner → reconciliation → missed occurrences.

Defend the hardest decision

Define misfire behavior: skip missed times, run the latest once, or catch up every occurrence within a cap. A recurring rule in a local timezone can encounter a missing or repeated wall-clock time; document the chosen behavior and persist computed UTC occurrence instants. Editing a schedule increments its version so queued old occurrences can be rejected or honored according to policy.

Failure and recovery analysis

A scanner creates a run then crashes before enqueueing it. The outbox repairs publication. A worker loses its lease while still calling the target; fencing protects internal state but cannot retract an external side effect. Require a target idempotency key derived from run ID, or document the remaining duplicate window. A heartbeat is not proof that a task succeeded.

Security and privacy boundary

Authorize target endpoints, validate payload size, and prevent scheduler use as an SSRF or privilege-escalation service.

Interview follow-ups with reasoning

Question: How do you cancel a running job?

Show answer and explanation

Answer: Record cancellation intent and cooperate with workers; external effects may need compensation.

Question: How do you recover a scanner partition?

Show answer and explanation

Answer: Resume due-time scanning and rely on unique occurrence insertion.

Question: Which metrics matter?

Show answer and explanation

Answer: Schedule-to-start lag, oldest due occurrence, lease churn and per-tenant starvation.

Operate and verify the design

Scheduled-to-start lag, oldest overdue occurrence, lease churn and tenant starvation.

Restart the scanner after occurrence commit but before enqueue and verify one durable run identity.

A second scenario to test transfer

A daily 09:00 job is scheduled in a local timezone that changes clocks. Persist the timezone and recurrence intent, then define whether a skipped local time runs at the next valid instant and how an ambiguous repeated time is handled. The next occurrence is calculated according to that documented policy, not a fixed 24-hour timer.

The scheduler publishes a due run and crashes before updating its row. How should recovery behave?

Show answer and explanation

Answer: Recompute or republish using the same occurrence identity. A unique run key prevents a second logical occurrence, and the worker applies the effect idempotently. Mark dispatch only after durable handoff, but assume the handoff may repeat. State whether delayed work is replayed, skipped, or coalesced after a prolonged outage.

Compare alternatives

ChoiceBenefitCost
Pre-materialized runsFast due-time scanHorizon maintenance
On-demand recurrenceLess stored future workMore scheduler computation
Catch-up policyPredictable outage recoveryMay skip or delay promised work
Fencing + idempotencySafe retriesExternal target support

A design-changing exercise

The task sent an email but crashed before marking success. Can exactly-once scheduling prevent a second email?

Show answer and explanation

Answer: Not by itself. The email boundary needs idempotency or reconciliation; the scheduler cannot atomically commit an unrelated provider effect.

Design workshop: materialize occurrences and lease attempts

Core scope is one-time/recurring schedules, bounded dispatch, retry and cancellation. Arbitrary code execution is a separate worker capability. Choose dispatch within one minute for admitted work, at-least-once task delivery and one logical occurrence per schedule/version/time. External exactly-once effects require target support.

Choose a rolling future occurrence horizon plus a due-time database index. next_run_at advances transactionally with unique occurrence creation. A scan claims bounded due batches using row locks/SKIP LOCKED where appropriate; this is queue-style work distribution, not a globally ordered read guarantee. Set runs.next_eligible_at for first dispatch and subsequent retry. Keep attempts(run,generation,worker,start,end,outcome) separate from the logical run.

sql
BEGIN;
SELECT id FROM runs
WHERE state='ready' AND next_eligible_at<=:now
ORDER BY next_eligible_at,id
LIMIT 100 FOR UPDATE SKIP LOCKED;
-- For each selected run: state=leased, generation+=1,
-- lease_until=now+lease; insert attempts and dispatch outbox.
COMMIT;

A worker commits result only if run generation and active lease/owner policy still match. Expired ownership can create duplicate execution, but cannot create a second logical occurrence. Send the stable run ID as the target idempotency key; using attempt generation as the external key would turn a retry into a new side effect. Use capped exponential backoff with jitter, for example min(300s,2^attempt seconds), maximum attempts and a DLQ with run/attempt evidence.

Logical run survives scanner and worker crashes. Scanner to Run DB/outbox: Insert occurrence and run R; unique schedule/version/time; Run DB/outbox to Executor: Publish R, possibly more than once; Executor to Run DB/outbox: Claim generation 4 with bounded lease; Executor to Task target: Invoke with stable target idempotency key R; Executor to Run DB/outbox: Lost response -> attempt unknown; reconcile; Executor to Run DB/outbox: Replacement generation 5 fences old internal completion
Scroll to inspect the diagram, or open it at full size.

Figure — Logical run survives scanner and worker crashes.

Define calendar behavior explicitly. For America/New_York daily 02:30, this scenario shifts a skipped spring-forward occurrence to the first valid local time after the gap, and runs a repeated fall-back local time once using the earlier offset. A daily 01:30 fall-back occurrence therefore runs once, not twice. Persist timezone database/version and resolved UTC occurrence identity. Different business policies are valid but need different test expectations; adding a fixed 24 hours loses calendar intent.

Chosen scheduler calendar and editing policy. Spring gap skips 02:30 / Run at first valid local time after gap / One resolved occurrence for local date; Fall repeats 01:30 / Earlier offset only / One occurrence, not two; Outage misses 10 daily runs / Coalesce under this schedule's policy / One catch-up run with skipped count recorded; Schedule edited / New version applies only to unclaimed future occurrences / Already leased work follows explicit cancellation policy; Cancel races with external call / Revoke generation and request stop / Cannot undo a target effect already accepted
Scroll to inspect the diagram, or open it at full size.

Figure — Chosen scheduler calendar and editing policy.

A tenant fair scheduler rotates eligible tenants and caps each claim batch/concurrency. At 100,000 midnight runs and a 60-second dispatch target, capacity needs about 1,667/s, separate from task duration. A one-hour occurrence horizon is insufficient for a long outage unless materialization catches up under the declared misfire policy. Query schedule version and enabled state before claiming; mark superseded pending occurrences canceled.

On lease expiry, preserve unknown attempt evidence and create a replacement only under retry policy. A target without idempotency may run twice; surface that limitation in schedule creation and operations. Monitor due lag, claim contention, attempt unknown age, per-tenant capacity and DLQ growth.

Exercise: Generation 4 sent an email but lost its response. Generation 5 invokes the same run. Does fencing alone prevent a duplicate email?

Show answer and explanation

Answer: No. Fencing protects scheduler state, not an external provider. The stable target key/status reconciliation must resolve the email operation; otherwise the delivery contract includes a duplicate window.

PostgreSQL explicit locks documents the transaction primitive; calendar and target-effect policies are application choices.

Technical references

Provider idempotency contract example.

PostgreSQL transaction isolation and concurrent updates.

Your study notes