You have one quarter to delete the shared staging environment that eight teams deploy into. Sequence the work so each step is reversible, say where the point of no return actually is, and name what cannot move.
Show the full answer Hide the answer
The sequence
- Inventory the claims, for four weeks. Record every defect staging actually caught and classify it: single-service regression, cross-service contract break, data-shape or data-volume problem, or operational (configuration, migration, startup order). Nothing is deleted, so this step is free and fully reversible. An environment is justified by the claims it proves, and you cannot relocate a claim you have not written down.
- Relocate single-service regressions to per-pull-request environments or a local container stack. Reversible: the shared environment is untouched.
- Relocate cross-service contract breaks to consumer-driven contract tests, published by each consumer and verified in the provider's pipeline. This is the long pole. Eight teams with two or three integrations each is on the order of 20 contracts to write and own.
- Relocate operational claims to a production canary with automated gates, plus a migration rehearsal against a restored snapshot of production-shaped data.
- Stop gating releases on staging while leaving it running. Still reversible in one configuration change: re-add the gate.
- Freeze staging deploys for two weeks and watch escaped-defect counts and change failure rate. If neither moves, delete.
Where data can diverge, and how you would know
The signal to watch through steps 5 and 6 is escaped defects by class, compared against the step-1 inventory. If contract breaks start reaching production, step 3 is incomplete and you reverse step 5 rather than pressing on. Aggregate change failure rate is too coarse: it moves late and does not say which claim you dropped.
The point of no return
Not the deletion. The deletion is reversible for as long as you keep the environment's infrastructure code and one dataset snapshot, which costs storage and nothing else. Keep both for a quarter.
The real point of no return is earlier and social: once teams stop maintaining their staging deployment manifests, re-standing the environment up costs weeks rather than hours. That happens somewhere around step 5, before anything is deleted, which is why step 5 is where you insist on the measurement.
The rollback at each stage
Steps 1 to 4 add capability and remove nothing, so rollback is abandonment. Step 5 is one gate configuration. Step 6 is a redeploy from retained infrastructure code, plus whatever manifest rot has set in, which is the cost the previous section names.
How long it really takes
Contract tests dominate, and they are not a platform team's work to do: each consumer writes its own. A quarter is plausible only if step 1 shows the cross-service claim is thin. If staging is catching contract breaks weekly, budget two quarters and say so at the start.
What cannot move
- Anything needing production-shaped data volume. A migration against 40 million rows behaves nothing like the same migration against 4,000, and an ephemeral environment with seeded data cannot prove it. That claim moves to a restored snapshot, not to a smaller environment.
- Anything depending on a third-party sandbox that exists once. A single vendor test account cannot be handed to forty ephemeral environments, so it stays behind a booking queue somewhere.
When this is the wrong answer
If step 1 shows staging catches cross-service breaks constantly and no team has contract tests, deleting it inside a quarter is a schedule, not a plan. The correct answer then is to keep the environment and spend the quarter on step 3, because the environment is a symptom of missing contracts, and removing a symptom first is how a migration turns into an incident.