advanced 3 min answer

Nine teams share one integration environment where all 34 services run together. The 900-test suite takes three hours and a red run is usually somebody else's deploy rather than a defect. You have been asked to move to per-service tests with virtualised dependencies without interrupting the fortnightly release train. What is the sequence, and where can it go wrong?

integration-test-boundariesshared-environmentmigrationtest-doublescan-i-deploy
Show the full answer Hide the answer

The sequence

  1. Classify the 900 tests by what they actually exercise. Instrument a full run and record which services' code each test executes. Expect a large share to touch exactly one service. Those move immediately at zero risk, because the shared environment was never providing them anything. This step is unglamorous and it usually delivers half the benefit.
  2. Build the new boundary for one service, chosen for being busy rather than important. Real database and real broker in containers; HTTP dependencies replaced by stubs built from recorded interactions. Reversible: nothing has been switched off.
  3. Run both suites in parallel for four to six weeks, with the shared environment still authoritative. Classify every disagreement individually. A test green in the new boundary and red in the old is either a shared-environment flake or a fake that lies, and those two have opposite implications. Averaging them or declaring a pass rate destroys the only information this stage produces.
  4. Replace what the shared environment was silently providing. It was the one place that version skew across services was ever observed. That capability has to be rebuilt explicitly as a deployment compatibility check - is this version verified against the versions currently deployed in the target environment - plus a thin post-deploy smoke test against production. Skipping this step is the single most common way the migration produces a worse system.
  5. Shrink the shared environment to the residue. The genuinely cross-service journeys, target 20 to 40 tests, and take the shared environment off the release gate so a red run no longer blocks nine teams.
  6. Point of no return: decommissioning. Do it only after the environment has been non-authoritative for a full release cycle with no escape it would have caught.

Where it diverges and how you would know

The recorded fixtures go stale, and a stale fixture passes. The safety of the whole scheme rests on one scheduled job: replay the recorded request set against the real providers nightly, diff the response shapes and status codes, and fail loudly on divergence. Teams build the virtualisation and skip the replay, which converts a three-hour honest suite into a twenty-minute dishonest one.

The second divergence is subtler. Stubs are written from the interactions the tests happen to make, so the error and limit behaviour of the real dependency - 429s, partial responses, 30-second timeouts - is usually absent. Assert on those explicitly or the new suite will be systematically optimistic about exactly the conditions that cause incidents.

Rollback

Every stage before 6 rolls back by reverting the release gate to the shared environment, which is why it stays running and authoritative for so long. The cost of keeping it an extra quarter is far below the cost of discovering in week three that step 4 was not done.

How long it really takes

Two to three quarters for 34 services, and the tooling is not the long pole - steps 1 and 3 are. Budget for classification and disagreement triage, not for containers.

When not to run this migration

With six services and one team, the shared environment is cheap and the coordination cost this migration removes does not exist. The thing being bought is independence between teams, so the return scales with the number of teams sharing the environment, not with the number of services. Below roughly three teams, keep the shared environment and spend the effort on making its data deterministic.