Test Double
also called Stub, Mock, Fake, Spy
Any stand-in for a real collaborator in a test - and the five kinds differ in what they let you assert, which is why using the wrong one produces tests that break on refactoring or pass while the code is wrong.
Two tests cover the same payment code. One breaks on every refactoring although behaviour is unchanged. The other passed for a year while the integration was broken in production. Both use "a mock", and both are using a kind of test double for the wrong purpose.
A test double is any object substituted for a real collaborator so a test can run fast and deterministically. The term is Gerard Meszaros's, from xUnit Test Patterns (2007), and the taxonomy matters because the choice of double determines what the test can assert, and therefore what it will and will not catch.
- Dummy — passed to satisfy a signature, never used. A null-ish placeholder.
- Stub — returns canned answers.
getPrice()always returns 100. Lets you set up a state and assert on the outcome. - Spy — a stub that also records what happened, so the test can check afterwards that a method was called.
- Mock — pre-programmed with expectations about the calls it should receive, and fails the test if they do not occur. Assertion lives in the double.
- Fake — a real, working, simplified implementation. An in-memory repository with genuine logic, a SQLite stand-in for Postgres.
Why it matters
The distinction that does the work is state verification versus interaction verification. A stub lets you assert on the result: given a price of 100, the total is 120. A mock asserts on the conversation: charge() was called once with these arguments.
Interaction verification couples the test to the implementation. It pins how the code achieves the result, so any refactoring that changes the call sequence breaks the test although behaviour is unchanged. That is the first test above, and it is why a mock-heavy suite makes a codebase expensive to improve: the tests become a cost of change rather than a protection against regression.
The second failure is subtler. A stub returning a hardcoded success verifies arithmetic on a fixture. It tells you nothing about whether the real collaborator behaves that way, and nothing in the suite can notice when it stops doing so. That is the second test.
Implementation patterns
- Prefer fakes for anything you own and use a lot. An in-memory repository with real behaviour — enforcing uniqueness, rejecting invalid input — gives fast tests that exercise logic and survive refactoring. Maintaining it costs perhaps a day and is repaid across hundreds of tests.
- Use stubs by default, mocks rarely. Assert on the outcome; reach for interaction verification only when the interaction is the behaviour — an email sent, an audit record written, a payment not charged twice.
- Put doubles at process boundaries you do not control. A double for a third-party HTTP API earns its place; one for a class in the same deployable pins your own design and verifies nothing external.
- Generate stubs from a contract, not by hand. A double derived from an OpenAPI schema or a verified consumer contract cannot drift from the shape the provider actually serves — which is the mechanism, rather than diligence, that keeps it honest.
- Double the failure modes. Timeout, 429, 5xx, malformed body, slow response — about 5 fixtures per external dependency at roughly 30 minutes each. Most libraries make the happy path trivial and the failures manual, so error-handling code is never executed in any test.
Industry example
Martin Fowler's "Mocks Aren't Stubs" (2007) is the reference discussion, and the distinction it drew — classical state-based against mockist interaction-based testing — settled into a default preference for state verification, with interaction verification reserved for genuine side effects. It shows in how the frameworks evolved: libraries built around strict expectations now default to lenient stubbing, because the accumulated experience was that over-specified interactions make refactoring expensive without catching more bugs.
Failure scenarios
- Over-mocked suites that break on refactoring, so behaviour-preserving improvements carry a test-rewriting cost and stop happening.
- Mocks of your own code, which produce tests of mocks calling mocks and verify nothing about the system.
- Silent drift. A hand-written stub of a third party ages into an accurate description of last year's API, and the suite stays green while telling you nothing.
- Happy-path-only doubles, leaving all error handling untested.
- Fakes that diverge from the real thing. An in-memory repository permitting what the database forbids, so tests pass and production rejects. Run one suite against both to pin them together.
- Assertions on call counts for calls that are implementation detail, producing failures nobody can interpret.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Stub | Fast, simple, asserts on outcomes, survives refactoring | Says nothing about whether the real collaborator behaves that way |
| Mock | Verifies side effects actually happened | Couples the test to the call sequence; breaks on refactoring |
| Fake | Real logic at test speed; robust to refactoring | Must be built and kept faithful to the real implementation |
When not to use it
When the integration itself is under test, use the real collaborator. A test of your database query belongs against a real database — a container starts in seconds — because a fake repository reproduces neither the query planner, the constraint, nor the transaction semantics, which is the whole subject of the test.
And do not double what is cheap and deterministic. A pure function, a value object, a local formatter: call it. A double adds indirection and a second thing to maintain for no speed gain.
The decision rule: double across a boundary that is slow, non-deterministic, costly or outside your control. Everything inside that boundary, use the real thing. A suite whose doubles sit at process boundaries and whose assertions are on outcomes survives refactoring and still catches regressions, which is the combination the taxonomy exists to make reachable.
Interview question
Q: A team's suite breaks on nearly every refactoring, and separately missed a production bug in an integration it has 100 tests for. Diagnose both.
What a strong answer covers: naming interaction verification as the cause of the first — mocks pinning call sequences rather than behaviour — and prescribing stubs with outcome assertions plus fakes for owned collaborators · naming the second as a double with nothing verifying it against reality, and distinguishing that from "mock less" · placing doubles at process boundaries and real objects inside them · generating stubs from a contract so drift is structurally prevented · and noting the two problems have opposite remedies, so one "use fewer mocks" mandate would fix one and worsen the other.
Quick check
Quiz: Which test double makes the test fail if an expected call does not happen? — A mock. The assertion lives in the double, which is also why it couples the test to the implementation.
Flashcard: When is interaction verification the right choice? — When the interaction is the behaviour: an email was sent, an audit record written, a payment not charged twice. For everything else, assert on the outcome with a stub or a fake.