A travel platform of the kind Expedia operates in — dozens of supplier APIs, each rate-limited and slow — runs its test suite against recorded stubs of every supplier. The suite is fast and green. What happens over the following year, and how do you know before a customer does?
Show the full answer Hide the answer
Second by second, what happens
Not seconds — months. That is the point of this failure: it has no incident, no alert and no moment.
Month 1. The stubs are recordings of real supplier responses. They are accurate. The suite is fast, deterministic and green, and the team rightly congratulates itself.
Month 3. A supplier adds an optional field and begins returning null where it previously omitted a key. The stub still returns the old shape. The code that would break on a null is never exercised. The suite remains green and has silently stopped testing that supplier.
Month 6. A supplier changes an error response from HTTP 400 with a body to HTTP 422 with a different body. The error-handling branch in the integration code is now dead in tests and wrong in production. Nobody has touched that code, so nobody looks.
Month 9. A supplier tightens rate limits and starts returning 429s under load. The stub has never returned a 429 in its life. There is no backoff code, because no test ever demanded any.
Month 12. The suite is green, it runs in four minutes, and it is now a high-fidelity test of what the suppliers were doing last year. Its greenness is evidence of nothing, and it is believed completely, because it was accurate when it was built and nothing announced that it had stopped being.
Where it amplifies
The confidence is the amplifier. A slow, flaky suite gets distrusted and supplemented. A fast, green, stale suite displaces the manual checking that would have caught the drift, so the better the virtualisation, the more thoroughly it removes its own safety net.
It amplifies a second way: stubs are usually recorded once, by whoever did the integration, and the recording is then edited by hand as tests need new cases. Within a year a meaningful fraction of the fixtures are hand-written approximations of a supplier's behaviour that the supplier never produced, and some of them were wrong on the day they were written.
What the user sees
A booking that fails for one supplier, in one circumstance, with an unhandled exception rather than a graceful fallback — the exact code path the stub guaranteed was tested.
What stops it
Mechanisms, not discipline:
- Contract verification against the real supplier, on a schedule. A nightly job that calls each supplier's sandbox or a low-cost real endpoint and asserts that the recorded fixtures still match the live shape. This is the single mechanism that makes virtualisation safe, and it is the one almost always omitted. It fails loudly, in its own job, without slowing the main suite.
- Re-record rather than hand-edit. Fixtures are generated artefacts with a recorded date. A fixture older than 90 days is a build warning.
- Stub the failures too. Every virtualised supplier needs fixtures for timeout, 429, 5xx, malformed body and slow response. Most stub libraries make the happy path easy and the failure modes manual, so the failure modes do not exist.
- A small real-integration suite on the critical path. Not every supplier, not every case: one booking per supplier per day against the real sandbox. Slow and flaky by nature, which is why it runs out-of-band and pages nobody — but it is the only thing that observes reality.
- Consume the supplier's change feed. Changelogs, deprecation headers, API version notices. Treat an incoming deprecation header as a build warning rather than a thing to read later.
When not to abandon the stubs
It is right, and this is not an argument against it. Dozens of rate-limited, slow, sometimes-charging third parties cannot be in the main test loop: the suite would take an hour, fail randomly, and cost money per run. The error is not using stubs; it is treating a recording as a contract. A stub is a performance optimisation over the real thing, and like any cache it needs an invalidation strategy. The nightly verification job is that strategy, and a team that has stubs without it has a cache with no TTL.