Four years ago a platform team adopted a cloud vendor's landing-zone reference architecture whole: hub-and-spoke network, three subscription tiers, a central firewall appliance, a shared CI account. The estate turned out to be three products and 40 engineers. The appliance adds a 12 ms hop to every internal call and caused two of the last five incidents. Auditors were told the architecture "conforms to the vendor reference". Sequence the migration to a simpler topology without losing the audit story.
Show the full answer Hide the answer
What went wrong is not the reference model
A reference model is a checklist of concerns and a vocabulary. It is written for the largest plausible adopter, because that is the only way one document serves everyone, and the adopter is expected to delete what does not apply. This team adopted it as a design instead of as a questionnaire, and the consequence is a topology sized for an estate they do not have.
The expensive part is the sentence given to the auditors. "Conforms to the vendor reference" made conformance the control, so every deletion now reads as a deviation from a control rather than as a design decision. That is why this is a migration with an evidence problem rather than a network change.
The artefact that has to exist before any change
A tailoring record: one row per element of the reference model, with the element, whether it is kept, replaced or deleted, the concern it was there to address, and how that concern is met now. The appliance's row says: kept concern is east-west traffic inspection and egress control; replaced by per-subnet security groups plus a managed egress gateway with flow logs to the same sink.
Write this before touching anything. It converts the auditor's unanswerable question — why do you differ from the reference? — into forty answerable ones, and it is the document that makes each following step defensible on its own.
The sequence, each step reversible
- Stand up the replacement egress path alongside the appliance. No traffic. Confirm flow logs arrive at the same destination with the same fields, because the audit trail is the thing that must not break.
- Generate both rule sets from one source. The appliance's rules and the security groups are rendered from a single definition, so the two paths cannot drift while both exist. This is where data diverges if you skip it, and drift is invisible until a rule exists on one path only.
- Move one product's east-west traffic by route-table change, lowest-traffic product first. Watch connection error rates and the 12 ms that should disappear from internal p99.
- Move the remaining two, one per week, each a single route-table change that reverts in minutes.
- Leave the appliance running with zero flows for 30 days. This is the rollback window and it is cheap.
- Decommission, and only now collapse the subscription tiers, because subscription changes are the genuinely hard-to-reverse part.
The point of no return
Not the route-table cutover — that reverts. It is step 6's subscription collapse, because identity assignments, policy scopes and billing hierarchies are re-created rather than moved. Everything before it is a change of path; that step is a change of container.
How long it really takes
The routing work is two to three weeks. The tailoring record and the conversation with the auditors is the long pole, usually a quarter, because it requires agreeing that conformance was never the control. Budget it that way at the start, or the network change will be blocked at the last gate by the one stakeholder who was not consulted first.
When not to do this at all
If a regulator or a customer contract names the vendor reference architecture as a requirement, the 12 ms and the two incidents are the price of the contract and the answer is to make the appliance highly available instead. Check that before writing the record. The other case for leaving it alone: if the estate is forecast to reach dozens of teams within a year, the topology is early rather than wrong, and migrating away and back costs more than the hop.