A B2B scheduling platform runs in one US region. 38% of revenue now comes from Australia and Europe, where p95 page load is 1.9 s against 420 ms in North America, and churn in those markets runs 2.3x the US rate. The team chooses one home region per tenant with a full stack in each, not global replication. Give the sequence under live traffic, say where data can diverge, and name the point of no return.
Show the full answer Hide the answer
The sequence
- Add a region attribute to every tenant and route on it, with every tenant set to
us-east. Nothing moves. The edge, the API gateway and every background job now read the attribute, so the routing layer is proven in production while there is exactly one possible answer. Reversible by deleting a column. - Separate global from regional data. Login identity, the tenant directory, billing and cross-tenant admin cannot be pinned - every region needs them. Move them behind an explicit global service with its own availability target before anything else, because discovering this coupling after a tenant has moved is the expensive order.
- Stand up the second region and run it empty, same pipeline, same configuration source, same alerts, for two weeks with no tenants on it. Most of the real cost surfaces here, in configuration drift and in jobs that assumed one region.
- Migrate one internal tenant, then one friendly customer. Per tenant: freeze writes (seconds to minutes for a scheduling product), copy, verify row counts and checksums, flip the region attribute, unfreeze. Per-tenant freeze is what makes this a sequence of small reversible steps rather than one cutover.
- Migrate in waves by timezone, during each wave's local night, keeping the source data for 30 days.
- Decommission the copied data only after a full business cycle - a month-end, a billing run, a support escalation - has passed in the new region.
Where data can diverge
- During the freeze window, if any writer bypasses the API: a cron job, an admin tool, a webhook handler, a support script. Inventory every writer before step 4; this is the usual source of silent loss.
- In the global tier, where a tenant's identity row says one region and its data lives in another. A single authoritative region attribute, read by everything and written by one process, is the only defence.
- In asynchronous work in flight: a queued email or report job enqueued before the flip and processed after it, pointing at the old database. Drain the tenant's queues inside the freeze.
- In caches and search indexes keyed without a region, which serve one region's data in the other until they expire.
Detection: a per-tenant reconciliation that compares row counts and a content checksum on the twenty highest-value tables, run hourly for the first week after each wave.
The point of no return
Not the flip, which is one attribute and reverses in seconds while both copies exist. The point of no return is deleting the source region's copy of a tenant's data, irreversible once the 30-day window closes. One threshold bites earlier: once tenants exist in two regions the global tier must serve both, so its availability target becomes the whole platform's ceiling, and returning to one region needs a second migration.
How long it really takes
Steps 1 to 3 are usually two thirds of the elapsed time and almost all of the surprises. For a platform of a few thousand tenants, plan two quarters with one squad, not one. Latency gains arrive only for migrated tenants, so publish the wave schedule to the markets that are churning.
When this is the wrong answer
If the latency problem is mostly front-end - uncompressed assets, no CDN, waterfall requests - a CDN and a round of asset work buys several hundred milliseconds in weeks, with no new region to operate. Measure server time against total page time before committing: if server time is 180 ms of a 1.9 s p95, the region is not the problem and this entire programme is misdirected. Equally, if the requirement turns out to be residency rather than latency, the design is different again: residency constrains which components may see the data, which this topology does not by itself guarantee.