Discord placed a Rust data-service layer in front of its message store partly to coalesce concurrent reads of the same row into a single database query. What does coalescing buy for tail latency and what does it pay for?
Show the full answer Hide the answer
What is gained
A hot partition's concurrency becomes bounded by the number of distinct keys rather than by the number of callers. When a very busy channel is opened by thousands of clients in the same second, the store sees one query for that row instead of thousands, so no queue forms at that partition and its p99 stays near its unloaded service time.
Discord's published account of its message storage (2023, alongside the move from Cassandra to ScyllaDB) describes the mechanism: requests are routed by consistent hash on the channel identifier so that all requests for a key reach the same process, the first request for a key starts a worker task, later requests for the same key subscribe to that task, and the single result is returned to every subscriber. The routing is what makes the coalescing possible — without key affinity there is nothing to coalesce.
The failure it removes is specific: unbounded concurrency against one partition, where each additional caller makes the partition slower, which makes callers wait longer, which raises concurrency further.
What is paid
- Shared latency. A subscriber that arrives 1 ms after the leader began a 400 ms query waits the remaining 399 ms, where on its own it might have been served from a faster path. Coalescing converts independent requests into one shared fate.
- Shared failure. The leader's timeout fails every subscriber at once. A 1-in-5000 error becomes 5000 errors in the same millisecond, and if every subscriber then retries, you have rebuilt the thundering herd you were preventing. Retry policy has to be written for the coalesced case.
- Staleness. The window in which requests join an in-flight query is a staleness window. For message bodies that is harmless; for presence or read-state it may not be.
- A new stateful hop. It has its own p99, its own deploys and its own routing table, and an unhealthy coalescer makes every key it owns slow at once. The blast radius follows the hash ring.
The rule for when it pays
Coalescing is worth a hop only when duplicate in-flight requests actually exist, and the expected number of them for a key is roughly (requests per second for that key) × (query latency in seconds):
- 20 rps on the hottest key with a 10 ms query: 0.2 duplicates. There is nothing to coalesce, and you have added a hop to every read.
- 5000 rps on the hottest key with a 10 ms query: 50 duplicates per query, so the store sees 2% of the reads it otherwise would.
Measure the per-key rate distribution first. If it is flat, this is machinery for a problem you do not have.
When not to build it
A cache with a single-flight lock around the miss path gets most of this effect without a new service, and it is where nearly every team should start: the same leader-and-subscribers logic, inside the process, with no routing layer to operate. Coalescing becomes worth its own tier when key affinity must be fleet-wide rather than per-instance, when the number of callers is large enough that per-instance single-flight still produces hundreds of duplicate queries, or when you want one place to enforce per-key concurrency limits.
It also does nothing for writes. Coalescing is a read-path tool; a hot partition's write contention needs bucketing, batching or a single-owner in-memory decider instead.