advanced 3 min answer

Roblox's January 2022 postmortem on the October 2021 incident states that the outage lasted 73 hours, that it ran its back-end services on a single Consul cluster, and that the single cluster supporting multiple workloads exacerbated the impact. Running one shared coordination cluster rather than one per workload is a real and defensible saving. What was that saving worth, and what decision rule follows?

robloxconsulblast-radiusconsolidationerror-budgetpostmortem
Show the full answer Hide the answer

The situation they were in

A coordination cluster is small, expensive to operate well, and shared by nature. Service discovery, configuration and secrets all want one authoritative view, and every additional cluster is another quorum to patch, upgrade, monitor and reason about. Roblox's own write-up says Nomad and Vault depended on Consul, so an unhealthy Consul meant the platform could neither schedule containers nor fetch production secrets.

What they chose, and why it fit

One cluster for everything. The saving is genuine: nodes, licences, upgrade effort, and the cognitive cost of a split registry. For a fast-growing platform, consolidating coordination is the choice most teams make and most would defend in review.

What it cost them

The postmortem reports 73 hours. Set that against an availability target. A 99.9% annual target allows about 8.8 hours of downtime a year; 73 hours is more than eight years of that budget spent in one event. At 99.99% it is more than seventy years of budget.

Now price both sides. The saving from consolidation is capped by the cost of the second cluster — a handful of nodes plus operational effort, on the order of tens of thousands of dollars a year at this scale. The exposure is not capped, because when the shared component is also the thing you need in order to diagnose it, recovery time is set by human investigation rather than by failover. Roblox's post notes that monitoring which would have given better visibility relied on affected systems, and that this hampered triage. That is the asymmetry: a bounded, known saving against an unbounded, unknown duration.

The decision rule

For each shared component, ask two questions. Does its failure stop every workload? Is recovery bounded by an automatic mechanism, or by somebody working out what is wrong? If the answers are yes and the latter, the duplicate is insurance priced in the tens of thousands against an event measured in days, and you buy it. If failure is partial, or recovery is a tested failover measured in minutes, consolidate and spend the money elsewhere, unless that component is also the only path to the diagnostics.

Roblox's own remediation list points the same way: remove circular dependencies in the observability stack, speed up bootstrapping, and move to multiple availability zones and data centres.

Where copying this would be a mistake

A team of four engineers should run one Consul cluster. Two clusters double the upgrade work and the failure surface they can actually attend to, and their realistic outage is hours, not days, because their estate fits in one person's head.

The transferable lesson is not "shard your service registry". It is that a single shared dependency converts a capped saving into an uncapped outage, and that the cheapest mitigations are usually not a second cluster at all: a static fallback for configuration, cached secrets with a long enough lease to survive a control-plane outage, and a diagnostic path that does not route through the thing being diagnosed.

Common weak answers

  • "Multi-region would have fixed it." The defect was lock contention and a storage-engine performance problem inside one cluster. Replicating the same software with the same load pattern replicates the bug.
  • "This proves consolidation is wrong." It proves consolidation must be priced against the worst credible duration rather than the average one.