A grocery marketplace in the mould of Instacart integrates with 140 retailer systems for inventory lookup and order injection. A platform engineer proposes building a virtual service for each one so the test suite can run without them. Roughly what does that cost per year to keep faithful and how many of the 140 are actually worth virtualising?
Show the full answer Hide the answer
The assumptions, stated
Three numbers decide this, and only one of them is usually known.
- Build cost per virtual service. Recording a real interaction corpus and shaping it into a stub that reproduces status codes, latency, rate-limit responses and error bodies: 2 to 4 engineer-days. Call it 3 days, or 24 hours.
- Rate of observable change per provider per year. A retailer's integration surface shifts when they change stock semantics, add a field, tighten a limit or alter an error code. For third-party enterprise systems, 1 to 2 changes a year is a defensible planning figure; some change quarterly and many never change.
- Re-sync cost per change. Diff the recorded traffic against the fixture, update it, re-run the affected tests: 4 to 8 hours, call it 6.
The arithmetic
Build: 140 × 24 h = 3,360 hours, about 1.8 engineer-years one-off.
Maintenance: 140 × 1.5 changes × 6 h = 1,260 hours a year, about 0.7 of an engineer permanently, before anyone has written a test.
The change rate dominates the error. At 0.5 changes per provider it is 0.25 of an engineer; at 4 it is 1.9. Nothing else in the estimate moves the answer as far, which means the first thing to measure is not the build cost but how often these 140 interfaces have actually changed in the last two years. That data is in the integration team's incident and ticket history.
What the number rules out
A permanent 0.7-engineer tax to maintain fakes for dependencies that most tests never touch is not defensible, and the failure mode is worse than the cost: an unmaintained fake does not fail, it passes. It keeps returning the shape the provider had in 2024 while the suite reports green, which manufactures confidence rather than removing it.
The decision rule: virtualise a dependency only where all three hold - the real endpoint is unavailable, rate-limited or has real-world side effects; the dependency sits on a path tested more than daily; and the interface is distinct rather than a copy of a shared shape. For 140 retailers that typically selects 10 to 20, which is roughly a month of build and a few weeks a year of upkeep.
The remaining 120 go behind the same adapter interface and share one generic stub that conforms to the internal contract rather than to each retailer's dialect. Their real behaviour is then verified by a scheduled replay: send the recorded request set to the live endpoint nightly or weekly, diff the response shape, and open a ticket on divergence. That catches drift for the whole tail at a fixed cost instead of a per-provider one.
When not to virtualise at all
If the providers publish sandboxes that are free, stable and side-effect-free, use them and build nothing - a sandbox that the provider maintains is a fake somebody else pays to keep faithful. And a team with fewer than about ten integrations should hand-write stubs per test rather than stand up virtualisation infrastructure, because the infrastructure only pays back when the fixture corpus is shared across many suites.
What a strong answer adds
The number that should be reported alongside the cost is fixture age: the median days since each virtual service was last reconciled against its real provider. A team that cannot produce that number does not know what its tests prove, and the metric is what turns drift from an invisible risk into a backlog item.