practice

Move Group

also called Migration Affinity Group, Move Wave Affinity Set

The set of workloads that must cross to the new infrastructure in the same wave because the round trips between them sit on a user-facing critical path, derived from observed network flows rather than from the application inventory.

rehostdata centre exitnetwork flowslatencywaves

A data-centre exit picks its first wave the way every exit does: the machines with the lowest CPU, because they look safest. Forty of them move over a weekend and on Monday three business applications are unusable. Nothing failed, nothing errored, and the CPU graphs on both sides look healthier than before.

The machines that looked idle were idle because they were waiting. Each one was half of a pair that exchanged hundreds of round trips per user request across a rack, at 0.1 to 0.5 ms a trip. Those trips now cross a link at 10 to 40 ms. An application making 200 sequential calls per page render pays 200 × 25 ms, which adds five seconds to a request that used to take 100 ms. A move group is the unit that prevents this: the set of machines whose mutual chatter means they travel together or not at all.

Why it matters

Chattiness is invisible in every artefact an exit programme normally has. It is not in the application inventory, not in the CMDB, not in the CPU and memory sizing that produced the target cost model, and not in the ownership map that produced the wave plan. It shows up only in flow data, and it is the single most common cause of a rehost wave that has to be rolled back.

The second reason is schedule. A discovered dependency in month eight, when the lease expires in month nine, is not an engineering problem any more. Move groups are how a programme converts an unknown into a fixed list early, and the list is what makes a wave plan honest.

Implementation patterns

  • Collect flow records for at least 30 days, from virtual switch counters, netflow or sFlow at the aggregation layer, plus sampled connection tables during the business peak. Thirty days is not padding: the month-end path is the dependency nobody documented.
  • Build the group as a transitive closure. Start from one machine, pull in every peer above a chosen threshold of connections or bytes, repeat until the set stops growing. Real estates produce groups of 5 to 40 machines that rarely match the application inventory.
  • Tune the threshold by what it produces. Too low and the closure swallows the whole estate through shared infrastructure; too high and it breaks pairs that matter. Exclude backup, monitoring and identity traffic explicitly, then re-run.
  • Order groups by criticality, but never define them by it. Criticality decides which group goes first; flows decide who is in it.
  • Publish the residual cross-group flows as the wave's known latency exposure, with the expected added round trips per request, so the risk is a number rather than a surprise.
  • Re-run the closure after each wave, since the previous wave changed the topology.

Industry example

Large data-centre exit programmes from roughly 2015 onward converged on this technique, and the major cloud migration tool families all ship dependency-mapping agents whose main output is exactly this closure. The archetype: an insurer exiting a leased facility inventoried 1,100 virtual machines across 90 "applications" and found that flow analysis produced 140 move groups, of which 11 spanned four or more application owners. The shared licence server and the reporting database appeared in more than half the groups, and no owner-based wave plan could have kept them with their callers.

Failure scenarios

  • CPU-based wave selection, which systematically picks the waiting half of a chatty pair.
  • Thirty days of data collected over a quiet month, missing the quarter-end path entirely.
  • The closure run once and then frozen while the estate keeps changing for six months.
  • Groups silently split to fit a weekend's change window, which reintroduces the exact failure the analysis existed to prevent; if a group does not fit the window, the window is wrong.
  • Shared services left ungrouped because they belong to nobody, so they move last and every earlier wave is degraded until they arrive.

Trade-offs

Choose Gains Pays
Flow-derived move groups Waves that do not break applications; a known latency exposure per wave Two to six weeks of collection and analysis before the first move
Owner-based waves Clean accountability; each team owns its wave Shared services split from their callers by construction
Criticality-based waves Risk ordering that a steering group understands Says nothing about who can be separated from whom

The analysis also produces an uncomfortable by-product: a picture of coupling that nobody asked for, including services that two teams each believed they owned alone.

When not to use it

An estate of 30 machines does not need flow analysis. Two engineers who know the systems will draw the dependencies on a whiteboard in an afternoon and be more accurate than a tool, and the collection alone costs longer than the whole exit's planning. The practice starts paying at roughly 100 machines, when no single person has seen the call graph.

It also stops being the deciding factor when chattiness is not on the critical path: an estate that communicates asynchronously over queues, or one already built as HTTP services with measured latency budgets and headroom, can be waved by owner. And if a dedicated sub-2 ms link runs for the whole exit, groups can be split deliberately, with the link's cost as the constraint instead.

Interview question

Q: You are nine months from a data-centre lease expiry with 600 virtual machines to move. The first wave broke three applications. How do you plan the remaining waves, and what would you tell the steering group about the extra three weeks you now want?

What a strong answer covers: that the deciding property is round-trip latency between components that talk many times per request, with the arithmetic to show it; deriving groups as a transitive closure from 30 days of flow data rather than from the inventory; the treatment of shared services; ordering groups by criticality while defining them by flows; and framing the three weeks as buying a fixed list of unknowns before the schedule can absorb any more surprises.

Quick check

Quiz: Why does choosing the least-busy machines for the first wave break applications? Because a machine is often idle precisely because it is waiting on a chatty partner, and separating the pair converts sub-millisecond round trips into tens of milliseconds each.

Flashcard: What is a move group and how is it built? The transitive closure of network flows above a threshold, computed from 30 days of flow records; it is the set of machines that must cross the link in the same wave, and it typically contains 5 to 40 machines that do not match the application inventory.