Roblox's published account of its October 2021 outage - 73 hours from 28 to 31 October - says the difficulty of diagnosing two primarily unrelated problems buried deep in its Consul implementation was largely responsible for the length. Write-ups of that postmortem describe the remediation order: a node replaced on a degraded-hardware theory, then the whole cluster doubled from 64 to 128 cores with faster storage on a traffic theory, then a restore from an earlier healthy snapshot, then a return to 64-core hosts. What did that sequence assume, and what should change once the second remediation fails?
Show the full answer Hide the answer
The situation they were in
One Consul cluster provided service discovery and configuration for every backend service. Roblox's report names two causes: a relatively new streaming feature which, under unusually high read and write load, produced excessive contention, and a pathological performance problem in BoltDB, the embedded store Consul uses for its write-ahead log, triggered by their particular load. Latency on the affected operations had typically sat under 300 ms and was now around 2 seconds. The remediation order comes from published readings of that report rather than from a timeline Roblox itself laid out step by step, so treat the sequence as the shape of the response rather than as a minute-by-minute record.
What the hypothesis sequence assumed
One cause, and a cause in a layer the responders owned. Hardware, then traffic, then state: three plausible single explanations, each about something the team controlled and could change.
That assumption is the expensive part. A system with two independent defects produces evidence that refutes every single-cause hypothesis while remaining fully consistent with "something we control is wrong", so the search stays inside one family and converges on nothing. Each refutation feels like progress and narrows nothing, because the real answer was a conjunction two layers below the team's own abstraction.
What each attempt cost beyond the time
Every remediation under an unconfirmed hypothesis is an experiment, and three of these four changed the system's state. Doubling the cores changed the performance envelope, so measurements taken before it were no longer comparable with measurements after. Restoring a snapshot replaced the state that held the evidence. Reverting to 64 cores added a third configuration to reason about. By the time a correct hypothesis was available, the system under test was not the system that broke, and evidence — not capacity and not effort — is the scarce resource in a long outage.
What to change after the second failure
- Write the hypothesis down before acting, with the observation that would refute it. An untestable remediation is a guess with a deployment attached.
- Change family deliberately. Name the families out loud — load, state, configuration, dependency internals, clock and schedule, and the layer below your own abstraction — and require the third hypothesis to come from one the first two did not touch. Both Roblox causes sat in dependency internals, the family teams search last because they have no instrumentation there.
- Preserve state before any destructive step. Snapshot, copy logs off the hosts, and keep one node untouched as a control. One untouched node is worth more than the capacity it costs, because comparison is the fastest diagnostic method available and a uniform fleet removes it.
- Instrument below your own layer. The decisive signals were inside Consul and inside its embedded store. A team whose dashboards stop at the edge of its own services cannot see the family where this cause lived, which is also why the hypotheses never went there.
When this is the wrong lesson to draw
Do not generalise this into "never remediate before you understand". In most incidents the cause is the last change, the restore-first instinct is correct, and organisations that demand root cause before action have longer outages rather than shorter ones. The discipline described here starts at the point where the first two hypotheses have failed, and until then acting fast on the obvious is right.
Copying the rest would also be a mistake. The lesson is evidence discipline and instrumentation depth, not an argument for running your own coordination layer or for avoiding Consul — a four-engineer team will never diagnose a defect in a distributed store's write-ahead log, and their correct answer is a managed service with a vendor whose on-call can.