Recorded Interaction Corpus
also called Traffic-Recorded Fixtures, Captured Interaction Set
A set of real request and response pairs captured from a live dependency and used to build and continuously re-verify its stand-in, so the fake reproduces evidence rather than the author's assumptions.
A hand-written stub encodes the author's model of a dependency. That is the same model the production code was written against, so a test running against it can confirm an assumption but never challenge one. The bugs that reach production are the ones in the space between the model and the real system, and a fake built from the model cannot contain them by construction.
A recorded interaction corpus replaces the model with evidence. Requests and responses are captured from the live dependency - including the 429s, the partial responses, the 30-second timeout, the field that is sometimes a string and sometimes null, the error body that does not match the documentation - and the virtual service is generated from that set.
The corpus then has a second job, which is the one teams skip: it is replayed against the real provider on a schedule, and the diff is the drift alarm.
Why it matters
Service virtualisation fails in one characteristic way. An unmaintained fake does not break; it passes. It keeps returning the shape the provider had on the day it was written while the suite reports green, which manufactures confidence rather than removing it. The corpus is what makes that detectable, because a recording is a dated artefact that can be re-run against reality.
The cost side is real. At roughly 3 engineer-days to build a virtual service and 1 to 2 observable changes per third-party provider per year at about 6 hours a change, a hundred-odd integrations carry close to 0.7 of an engineer permanently in upkeep alone. A corpus with automated replay converts most of that from manual reconciliation into a failing scheduled job.
Implementation patterns
- Capture at the boundary, in production or in a pre-production environment carrying real traffic - a proxy, a sidecar or the client's own interceptor. Capturing from tests only re-records the assumptions.
- Scrub before it leaves. The corpus contains customer data and inherits its handling obligations until it does not. Tokenise identifiers consistently so referential relationships in the corpus survive scrubbing.
- Keep the awkward recordings deliberately. Sample by response class, not uniformly: every distinct status code, the slowest decile, and any response that failed to parse. A corpus of happy paths is a hand-written stub with extra steps.
- Stamp every corpus with a capture date and publish fixture age - the median days since each virtual service was reconciled. A team that cannot produce that number does not know what its tests prove.
- Replay nightly or weekly against the real provider, diffing status codes and response shape rather than exact bodies, and open a ticket on divergence. Shape-level diffing is what keeps the signal usable; body diffing produces noise from timestamps and identifiers.
- Share one generic stub for the tail. Where dozens of providers sit behind a common internal adapter, virtualise the distinct shapes and let the rest share a stub conforming to the internal contract, covered by replay rather than by fixtures.
Industry example
Marketplaces integrating with large numbers of retailer or logistics systems face this shape directly: a grocery platform in the mould of Instacart may carry well over a hundred inventory and order-injection integrations, most unavailable for testing on demand and several with real-world side effects if called. The workable structure is consistently the same - a corpus recorded from live traffic for the highest-volume and most distinct providers, a shared adapter stub for the long tail, and scheduled replay as the only thing standing between the suite and quiet obsolescence.
Failure scenarios
- The corpus ages silently because replay was never built, and the suite certifies an interface that changed a year ago.
- Scrubbing destroys the data's shape - identifiers randomised inconsistently, so relationships between recorded calls no longer hold and multi-step flows cannot be replayed.
- Only happy paths recorded, so the suite is systematically optimistic about exactly the conditions that cause incidents.
- Replay diffs on full bodies and produces daily false alarms, so it is muted, which is worse than not having it.
- The corpus becomes a data liability because scrubbing was deferred and a test environment now holds production records under weaker controls.
Trade-offs
The corpus buys fidelity you cannot write by hand, and it pays in a capture pipeline, a scrubbing step with real compliance weight, storage, and a scheduled job that will page someone. Hand-written stubs are free, immediate and wrong in ways nobody can see. A middle position that works: record the corpus for the handful of dependencies whose misbehaviour would cause an incident, hand-write the rest, and be explicit about which is which.
When not to use it
If the provider runs a free, stable, side-effect-free sandbox, use it and build nothing - a sandbox is a fake somebody else pays to keep faithful. If the dependency is internal and its team publishes a verified contract, that contract is a better artefact than a recording, because it is maintained by the people who change the thing. And a team with fewer than about ten integrations should hand-write stubs per test: the corpus infrastructure only pays back when fixtures are shared across many suites.
Interview question
Q: You want your integration tests to stop depending on a flaky third-party sandbox. Walk me through building a recorded corpus for it, and tell me what you would put in place so the corpus does not quietly rot.
What a strong answer covers: capture at the client boundary from real traffic rather than from tests · scrub with consistent tokenisation so multi-step flows survive · sample by response class so error and slow responses are represented · publish fixture age as a visible metric · scheduled replay against the live provider diffing shape and status rather than bodies · a ticket-generating divergence alarm · and the judgement that a stable vendor sandbox would make the whole exercise unnecessary.
Quick check
Quiz: What is the single job that separates a useful recorded corpus from a slowly rotting one? Scheduled replay against the real provider with a shape-level diff.
Flashcard: Why record a stub instead of writing one? - A written stub encodes the same assumption the production code encodes, so it can only confirm it; a recording carries the provider's real errors, latency and oddities, at the cost of scrubbing and a capture date.