A data-centre exit must move about 600 virtual machines in nine months. The first wave moved 40 machines chosen because they were the least busy and three business applications became unusably slow overnight. What should decide which machines travel in the same wave?
Show the full answer Hide the answer
The deciding property
The one fact that settles it is round-trip latency between components that talk to each other many times per request. Inside a rack, two machines are roughly 0.1 to 0.5 ms apart. Across a link between a data centre and a cloud region, they are on the order of 10 to 40 ms apart. An application that makes 200 sequential calls to its database per page render — an ordinary number for a chatty enterprise application built against a local database — pays 200 × 25 ms, which is five seconds added to a request that used to take 100 ms.
Nothing about that shows up in the CPU graph that picked the first wave. The least-busy machines are exactly the ones that are idle because they are waiting, and splitting them from their partners is how an application that nobody considered risky becomes unusable.
A move group is therefore the transitive closure of the flows above a threshold: start from one machine, pull in every peer it exchanges more than a chosen volume or connection count with, repeat until the set stops growing. Real estates produce groups of 5 to 40 machines, and the groups rarely match the application inventory.
How to get the data
Flow records over at least 30 days, so that month-end batch paths are included, from the virtual switch counters, netflow or sFlow on the aggregation layer, plus connection tables sampled during the peak. Thirty days matters: a dependency that fires once a month is exactly the one nobody documented, and it will hold the exit hostage in month eight.
Why the other options fail
- By application owner. This is the plan most exits start with, and it is right about accountability and wrong about topology. Shared services — the identity server, the file share, the licence server, the reporting database — belong to nobody and are in everyone's critical path, so owner-based waves split them by construction.
- By operating system version. A real efficiency for the rebuild tooling and a genuine cost saving on a rehost, but it optimises the thing that costs days while ignoring the thing that breaks applications. Use it to order work inside a move group.
- By CPU and memory size. This makes the target bill predictable, which matters to the business case, and it has no relationship to which machines can be separated.
- By criticality, least important first. Sound instinct for risk, and it is how the stem's team already failed: the low-criticality machines were the idle halves of chatty pairs. Criticality should order the groups, never define them.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| Applications are async or batched over a queue | Owner-based waves | Chattiness stops being on the critical path |
| A dedicated sub-2 ms link runs for the whole exit | Owner or size-based waves | The latency penalty disappears and the link cost is the constraint |
| The estate is already HTTP services with latency budgets | Owner-based waves | Call depth is known and already budgeted |
| The exit deadline is under three months | Largest-group-first | Discovery time now exceeds the value of optimal grouping |
When this is the wrong answer
An estate of 30 machines does not need flow analysis. Two people who know the systems can draw the dependencies on a whiteboard in an afternoon and be more accurate than a tool. The analysis starts paying at roughly 100 machines, when no individual has seen the whole call graph and the undocumented dependency is a certainty rather than a risk.