Slack Job Queue
Slack describes a Redis queue outage caused by backlog and memory pressure.
Slack describes a Redis queue outage caused by backlog and memory pressure. Its incremental redesign put Kafka before Redis rather than immediately replacing every queue behavior. Kafkagate accepted jobs into Kafka and relay components fed the existing execution layer. Preserving established scheduling and deduplication behavior made the migration smaller. This historical account illustrates durable buffering and controlled change, not unlimited downstream throughput. Source publication: 2017-12-06; updated 2020-06-25. The walkthrough below is an original interview exercise, not an undocumented claim about the company.
The mechanism at a glance
Figure — Application → Kafka ingress (submit); Kafka ingress → Durable log (durable append); Durable log → Relay (consume); Relay → Redis work queues (bounded transfer); Redis work queues → Workers (claim work); Workers → Redis work queues (finish or retry)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Calculate the overload window
In an independent queue exercise, arrival is 12,000 jobs/s and execution is 8,000/s. The difference adds 4,000 jobs/s to backlog. After 15 minutes that is 3.6 million jobs. A durable buffer buys recovery time but leaves the throughput deficit unchanged.
2. Bound the staging queue
A relay should stop moving work into a memory queue when it reaches a safe high-water mark. Resume below a lower threshold to avoid rapid oscillation. Include in-flight jobs and retries in capacity accounting; pending length alone understates memory use.
3. Preserve execution semantics
Moving the ingestion boundary does not automatically preserve delayed scheduling, priority or deduplication. Write down each old behavior, identify the new owner and test the transition. The worker must still make side effects safe under repeated delivery.
4. Plan rollback with offsets
For a generic migration, route a small cohort through the new ingress and compare accepted-to-completed counts. Track log offsets and relay acknowledgements. Rollback must not abandon jobs accepted only by the new path, so drain or replay them deliberately.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
job(id,type,arguments_hash,priority,not_before)
relay_checkpoint(partition,offset)
execution(job_id,attempt,state)
staging_budget(pending_bytes,inflight_bytes,limit)Worked example
Original exercise: a downstream database is slow for ten minutes. Pause relay transfer when staging memory reaches its limit, keep durable ingress within a finite retention budget, and calculate the drain time after recovery. If arrivals remain 12,000/s and execution recovers only to 12,000/s, backlog never shrinks.
Failure walkthrough
Original failure probe: relay enqueues a job but crashes before checkpointing the log offset. The same job will be transferred again. Stable job IDs and idempotent execution must tolerate it. Committing the offset before enqueue would instead create a job-loss window.
Figure — Accept into durable log → Relay under staging budget → Worker claims job → Commit business effect → Acknowledge completion safely
Decisions and trade-offs
| Decision | Useful when | Cost to explain |
|---|---|---|
| Adopt the mechanism | The same workload constraint is demonstrated | Validate with your own measurements |
| Keep a simpler design | Your scale and guarantees are already met | Monitor the trigger for changing it |
Check your understanding
After recovery, service reaches 16,000 jobs/s while arrival remains 12,000/s. How long to drain 3.6 million jobs?
Show answer and explanation
Answer: The net drain is 4,000 jobs/s, so the ideal lower bound is 900 seconds or 15 minutes. Retries, uneven job costs and other limits can extend it.
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Continue the connection
Study Job Scheduler and explain which guarantee from this lesson carries into that topic.
After recovery, service reaches 16,000 jobs/s while arrival remains 12,000/s. How long to drain 3.6 million jobs?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.