advanced 3 min answer

After its October 2021 outage Roblox said it was moving to multiple availability zones and data centres from a single-data-centre footprint of more than 18000 servers and 170000 containers coordinated by one shared cluster. Sequence that migration under live traffic and say where it can go wrong.

multi-regionrobloxservice-discoverymigrationblast-radius
Show the full answer Hide the answer

The sequence

Roblox's published writeup (January 2022) records the facts this plan has to work with: a 73-hour outage from 28 to 31 October 2021, a single data centre, a single coordination cluster serving many workloads which the company identified as having made the impact worse, and monitoring that depended on the systems that were failing. Each step below is reversible on its own.

  1. Split the coordination layer by criticality. The documented aggravating factor was one cluster serving everything. Separating the cluster that discovers production services from the one that serves configuration to batch jobs reduces risk before the topology changes, and it is the only step that pays off even if the rest is cancelled.
  2. Break the observability dependency. Telemetry, dashboards and the incident channel must not resolve through the layer being changed, because nothing later on this list is safe to attempt while the diagnostic path shares a failure domain with its subject.
  3. Make placement explicit. Label every workload with the zone it runs in and make the scheduler respect it. Nothing has moved yet; you have only made the concentration visible, which is reliably the moment somebody finds a stateful service with no replica anywhere.
  4. Replicate stateless capacity into the second zone and steer a small share of traffic — 1%, then 10% — with the ability to return it in one change. Measure cross-zone latency on the real call graph: a service making 30 internal calls per request pays the 1 to 2 ms round trip 30 times.
  5. Move stateful services one at a time, each with its own replication posture and its own drill. This is the long part and the part that decides whether the programme stays honest.
  6. Rehearse losing a zone before claiming the capability. Drain one deliberately at a low-traffic hour and keep the measurement.

Where data can diverge

Anything with a quorum. Stretching a coordination cluster across zones changes its own latency and failure behaviour: three nodes split one-one-one survive a zone loss, and three nodes split two-one lose quorum when the two-node zone is the one that fails — the configuration teams reach by accident when a region has only two usable zones. You would know from a continuously running read-your-writes probe against the coordination layer reported per zone, not from a health endpoint reporting the leader's opinion.

The point of no return

The first stateful promotion. Until a datastore's authoritative copy moves, every step is a traffic or label change. Once it moves, rollback requires a reverse replication path that must have been built before the cutover, not after it.

The rollback at each stage

Steps 1 to 4 roll back by changing a weight or a label. Step 5 rolls back only while the old primary is still receiving changes, which means dual-write or reverse replication for the whole window, and the window is the real cost of the migration, not the cutover.

How long it really takes

At this scale the stateless phase runs in quarters and the stateful phase in years. Roblox described the work as in progress in its January 2022 writeup, months after the event. The capability becomes real only when an unplanned zone event confirms it, which is why step 6 exists.

When this is the wrong programme

A platform whose incidents have all been software-caused does not get safer by adding zones. Zones address correlated physical failure. The two aggravating factors in this incident were a shared coordination cluster and an observability dependency, and steps 1 and 2 address both without touching the topology at all. Do those first, measure the next incident class, and be willing to stop there. Copying the zone programme because a large platform announced one, without checking that your failures are zonal, buys years of work against a risk you may not have.