Job Scheduler
Design a service that runs delayed and recurring jobs with retries, cancellation, and multiple worker regions.
Interview scope and guarantees
Create one-time or recurring jobs, pause schedules, trigger due occurrences, inspect outcomes and retry failed work. Define timezone and daylight-saving semantics. A schedule is not an execution: each occurrence has its own stable identity.
Capacity worksheet
Assume 10 million schedules with 1 million due per hour: 278/s average. If 100,000 are due at midnight, a one-minute dispatch target requires 1,667/s before execution cost. Partition by due-time bucket and tenant so one synchronized tenant cannot monopolize the fleet.
Concrete API contract
POST /schedules {expression,timeZone,payload,target}
PATCH /schedules/{id} {expectedVersion,enabled,expression}
GET /runs/{id}
POST /runs/{id}/retryData model and access paths
schedules(id PK,tenant_id,expression,time_zone,version,enabled,next_run_at,misfire_policy)
occurrences(schedule_id,schedule_version,scheduled_time,run_id); UNIQUE(schedule_id,schedule_version,scheduled_time)
runs(id PK,state,generation,next_eligible_at,lease_until); INDEX(state,next_eligible_at)
attempts(run_id,generation,worker_id,started_at,finished_at,outcome); PK(run_id,generation)
outbox(event_id PK,run_id,published_at)Evolve a solution and explain each change
Figure — Three architecture decisions for Job Scheduler, including the pressure each introduces.
Step 1: Persist due work
Store job definitions and next eligible execution time. A timer in one process disappears on restart.
Step 2: Lease executions
Claim due runs atomically with attempts, deadlines, and retry policy. Expired workers can finish after a replacement starts.
Step 3: Fence effects and calendars
Separate scheduled runs from attempts; define timezone, missed-run, and overlap rules. A successful task side effect can precede an ambiguous acknowledgement.
Responsibility overview
Figure — Connected responsibilities for Job Scheduler. Trace the authoritative and derived paths separately.
Worked end-to-end scenario
A daily job has run identity (job J, scheduled time T). Worker W1 leases attempt 1 and calls an external service. Its acknowledgement is lost and the lease expires. W2 retries attempt 2 using the same run-level idempotency identity for the external effect. W1’s later status update is fenced by its old attempt token. Scheduling tomorrow is distinct from retrying today’s run. For a daylight-saving transition, the schedule must say whether a repeated local time yields one run or two and how a nonexistent local time is handled. On a prolonged outage, choose skip, coalesce, or bounded catch-up explicitly.
Why these access paths matter
Index runnable executions by next eligible time and state. A unique job-plus-scheduled-time key prevents duplicate logical runs. Store attempts separately from run outcome and maintain a current fencing token. Partition due work into bounded claim batches; a lease is ownership of an attempt, not proof that the external task has never executed.
Build the baseline first
- Schedule API → durable rule → due-time index.
- Due scanner → unique occurrence and outbox.
- Queue → leased worker → durable outcome.
Evolve the design under load
- Time buckets → sharded scanners → fair queues.
- Worker pools → per-target concurrency budgets.
- Overdue scanner → reconciliation → missed occurrences.
Defend the hardest decision
Define misfire behavior: skip missed times, run the latest once, or catch up every occurrence within a cap. A recurring rule in a local timezone can encounter a missing or repeated wall-clock time; document the chosen behavior and persist computed UTC occurrence instants. Editing a schedule increments its version so queued old occurrences can be rejected or honored according to policy.
Failure and recovery analysis
A scanner creates a run then crashes before enqueueing it. The outbox repairs publication. A worker loses its lease while still calling the target; fencing protects internal state but cannot retract an external side effect. Require a target idempotency key derived from run ID, or document the remaining duplicate window. A heartbeat is not proof that a task succeeded.
Security and privacy boundary
Authorize target endpoints, validate payload size, and prevent scheduler use as an SSRF or privilege-escalation service.
Interview follow-ups with reasoning
Question: How do you cancel a running job?
Show answer and explanation
Answer: Record cancellation intent and cooperate with workers; external effects may need compensation.
Question: How do you recover a scanner partition?
Show answer and explanation
Answer: Resume due-time scanning and rely on unique occurrence insertion.
Question: Which metrics matter?
Show answer and explanation
Answer: Schedule-to-start lag, oldest due occurrence, lease churn and per-tenant starvation.
Operate and verify the design
Scheduled-to-start lag, oldest overdue occurrence, lease churn and tenant starvation.
Restart the scanner after occurrence commit but before enqueue and verify one durable run identity.
A second scenario to test transfer
A daily 09:00 job is scheduled in a local timezone that changes clocks. Persist the timezone and recurrence intent, then define whether a skipped local time runs at the next valid instant and how an ambiguous repeated time is handled. The next occurrence is calculated according to that documented policy, not a fixed 24-hour timer.
The scheduler publishes a due run and crashes before updating its row. How should recovery behave?
Show answer and explanation
Answer: Recompute or republish using the same occurrence identity. A unique run key prevents a second logical occurrence, and the worker applies the effect idempotently. Mark dispatch only after durable handoff, but assume the handoff may repeat. State whether delayed work is replayed, skipped, or coalesced after a prolonged outage.
Compare alternatives
| Choice | Benefit | Cost |
|---|---|---|
| Pre-materialized runs | Fast due-time scan | Horizon maintenance |
| On-demand recurrence | Less stored future work | More scheduler computation |
| Catch-up policy | Predictable outage recovery | May skip or delay promised work |
| Fencing + idempotency | Safe retries | External target support |
A design-changing exercise
The task sent an email but crashed before marking success. Can exactly-once scheduling prevent a second email?
Show answer and explanation
Answer: Not by itself. The email boundary needs idempotency or reconciliation; the scheduler cannot atomically commit an unrelated provider effect.
Design workshop: materialize occurrences and lease attempts
Core scope is one-time/recurring schedules, bounded dispatch, retry and cancellation. Arbitrary code execution is a separate worker capability. Choose dispatch within one minute for admitted work, at-least-once task delivery and one logical occurrence per schedule/version/time. External exactly-once effects require target support.
Choose a rolling future occurrence horizon plus a due-time database index. next_run_at advances transactionally with unique occurrence creation. A scan claims bounded due batches using row locks/SKIP LOCKED where appropriate; this is queue-style work distribution, not a globally ordered read guarantee. Set runs.next_eligible_at for first dispatch and subsequent retry. Keep attempts(run,generation,worker,start,end,outcome) separate from the logical run.
BEGIN;
SELECT id FROM runs
WHERE state='ready' AND next_eligible_at<=:now
ORDER BY next_eligible_at,id
LIMIT 100 FOR UPDATE SKIP LOCKED;
-- For each selected run: state=leased, generation+=1,
-- lease_until=now+lease; insert attempts and dispatch outbox.
COMMIT;A worker commits result only if run generation and active lease/owner policy still match. Expired ownership can create duplicate execution, but cannot create a second logical occurrence. Send the stable run ID as the target idempotency key; using attempt generation as the external key would turn a retry into a new side effect. Use capped exponential backoff with jitter, for example min(300s,2^attempt seconds), maximum attempts and a DLQ with run/attempt evidence.
Figure — Logical run survives scanner and worker crashes.
Define calendar behavior explicitly. For America/New_York daily 02:30, this scenario shifts a skipped spring-forward occurrence to the first valid local time after the gap, and runs a repeated fall-back local time once using the earlier offset. A daily 01:30 fall-back occurrence therefore runs once, not twice. Persist timezone database/version and resolved UTC occurrence identity. Different business policies are valid but need different test expectations; adding a fixed 24 hours loses calendar intent.
Figure — Chosen scheduler calendar and editing policy.
A tenant fair scheduler rotates eligible tenants and caps each claim batch/concurrency. At 100,000 midnight runs and a 60-second dispatch target, capacity needs about 1,667/s, separate from task duration. A one-hour occurrence horizon is insufficient for a long outage unless materialization catches up under the declared misfire policy. Query schedule version and enabled state before claiming; mark superseded pending occurrences canceled.
On lease expiry, preserve unknown attempt evidence and create a replacement only under retry policy. A target without idempotency may run twice; surface that limitation in schedule creation and operations. Monitor due lag, claim contention, attempt unknown age, per-tenant capacity and DLQ growth.
Exercise: Generation 4 sent an email but lost its response. Generation 5 invokes the same run. Does fencing alone prevent a duplicate email?
Show answer and explanation
Answer: No. Fencing protects scheduler state, not an external provider. The stable target key/status reconciliation must resolve the email operation; otherwise the delivery contract includes a duplicate window.
PostgreSQL explicit locks documents the transaction primitive; calendar and target-effect policies are application choices.