Discord Message Storage
Discord’s account describes moving its message store from Cassandra to ScyllaDB and adding Rust data services.
Discord’s account describes moving its message store from Cassandra to ScyllaDB and adding Rust data services. Messages are partitioned by channel and time bucket. Request coalescing and channel-based routing reduce duplicate concurrent reads before they reach storage. The migration included dual writes and comparison reads. The account makes clear that changing the database alone did not eliminate hot-partition concerns. Source publication: 2023-03-06. The walkthrough below is an original interview exercise, not an undocumented claim about the company.
The mechanism at a glance
Figure — API callers → Channel router (channel key); Channel router → Data service (stable route); Data service → Shared in-flight read (join or create task); Shared in-flight read → Message store (one read); Message store → Subscribers (fan out result); Subscribers → API callers (responses)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Choose the history access path
In a new chat service exercise, list the newest 50 messages in one channel. A partition key of channel plus bounded time bucket supports this query and limits partition growth. A global message ID table alone requires a secondary access path for channel history.
2. Separate coalescing from caching
Coalescing shares work only while an identical read is in flight. Caching reuses a completed result later. The former need not choose a long freshness interval, but requests must agree on authorization, projection and consistency requirements before sharing a result.
3. Apply backpressure before storage
Cap in-flight reads per channel and globally. If 10,000 users open one announcement, one shared fetch can satisfy many, but a stream of distinct reads still needs admission control. Use separate metrics for requests received, storage queries issued and callers waiting.
4. Migrate without assuming equality
For an independent migration, dual-write through a durable source or reconciliation process, backfill historical ranges, and compare sampled reads including deletes. Define the cutover and rollback authority. A passing count comparison is weaker than comparing message content, order and tombstones.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
messages(channel_id, time_bucket, message_id, body)
read_key(channel_id, bucket, cursor, limit, auth_scope)
inflight(read_key -> shared future)
migration(range_id, checkpoint, validation_status)Worked example
Original exercise: 1,000 authorized readers request the same page during a 40 ms storage read. With routing and a shared in-flight entry, the service can issue one read for that page. If every reader requests a different cursor, coalescing provides little benefit; concurrency limits must still protect storage.
Failure walkthrough
Original failure probe: the shared read times out. Complete every subscriber with a bounded error, remove the in-flight entry and limit retries. Leaving the failed future in the map permanently poisons that page. One caller cancelling must not cancel work still needed by other subscribers.
Figure — Readers request the same page → Route by channel → Join one in-flight task → Fetch storage once → Complete all authorized subscribers
Decisions and trade-offs
| Decision | Useful when | Cost to explain |
|---|---|---|
| Adopt the mechanism | The same workload constraint is demonstrated | Validate with your own measurements |
| Keep a simpler design | Your scale and guarantees are already met | Monitor the trigger for changing it |
Check your understanding
Why can hot partitions remain after adopting a faster database?
Show answer and explanation
Answer: A faster engine raises capacity but does not remove skew. One key can still exceed a node’s service budget. Routing, coalescing, admission control and a suitable partition model address different parts of that problem.
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Continue the connection
Study WhatsApp and explain which guarantee from this lesson carries into that topic.
Why can hot partitions remain after adopting a faster database?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.