intermediate 3 min answer

A grocery-delivery platform of the kind Instacart operates has 1,200 integration tests, all green, and an integration with the inventory service that has broken twice this quarter. The tests mock the inventory client. A developer proposes mocking less. Is that the fix?

integration testingmockstest boundariescontract testingdiagnosis
Show the full answer Hide the answer

The first three things I would look at

  1. What exactly broke, both times. Not "the integration" — the specific failure. A changed field name, a changed status code, a timeout, a semantic change in a value, a rate limit. The category determines the fix, and reaching for "mock less" before knowing it is a guess with a quarter's work attached.
  2. Who owns the mock and when it was last updated against reality. A hand-written mock reflects one developer's reading of the API on one day. If the inventory service has shipped 40 times since, the mock is testing a historical API.
  3. Whether the tests assert on the interaction or only on the outcome. Tests that verify "the client was called with these arguments" pin an interaction that may be wrong. Tests that assert only on the final outcome through a mock that returns a hardcoded success are verifying arithmetic on a fixture.

The diagnosis

The mock is not the problem; the absence of any verification that the mock matches reality is. These are different defects with different fixes, and conflating them is why "mock less" is the instinctive and usually wrong answer.

A mock is a statement about how a dependency behaves. Like any statement, it can become false, and nothing in the test suite is capable of noticing. The suite is internally consistent and green, and its greenness carries no information about the inventory service. That is true whether there are 1,200 tests or 12,000, so adding real integrations to some of them changes the number of tests that are meaningful without changing the mechanism that lets the rest rot.

Why "mock less" is the misleading fix

It sounds like increased realism and it buys less than it costs:

  • Real dependencies in the main suite make it slow and flaky, which leads to re-running, which destroys the signal from every other test in the suite. A quarter's gain in realism for a permanent loss in trust is a bad trade.
  • It does not scale past one dependency. Hitting the real inventory service means hitting whatever the inventory service depends on, and at some depth the suite requires the whole estate deployed and the tests fail for reasons unrelated to the change under test.
  • It still would not have caught either failure, unless the real service happened to be running the new version at the moment the test ran — which is a timing coincidence, not a mechanism.

The fix

Keep the mocks and add the thing that verifies them.

  1. Consumer-driven contract tests. Record what this service actually requires of the inventory service, and verify the inventory service's pipeline against that contract before its changes merge. This moves detection to the provider, before the break, and it is the structural answer. The two failures this quarter would both have been caught at the provider's merge.
  2. Fail on contract drift, not on a schedule. A nightly job comparing mock fixtures against the live service's schema catches drift even without the provider's cooperation. Weaker than contracts, and it works when the provider is a different company or an unwilling team.
  3. Mock the failures, not only the happy path. Both of this quarter's breakages likely had an error-handling dimension. If no test ever makes the inventory client time out or return 503, that code has never run.
  4. Keep a handful of real integration tests, out of band. Perhaps twenty, against the real service, running on a schedule rather than on every commit, paging nobody and reporting to the owning team. Slow and flaky is acceptable when it is out of the critical path, and this is the only component that observes reality.

When not to keep the mocks

If the suite's mocks are mocking your own code — internal classes, your own repository layer, modules in the same deployable — then mock less is exactly right and the problem is a different one. Mocks belong at process boundaries you do not control. A mock of an in-process collaborator pins your own design in place and makes refactoring expensive while verifying nothing about the outside world. The 1,200 tests are worth auditing for that, and it is a plausible second finding, but it is not what broke the inventory integration.