intermediate 4 min answer Multiple choice

Twelve services, six teams, one shared event schema plus a dozen REST APIs between them. Breaking changes reach other teams about twice a month, always discovered in the shared staging environment a day or two after the merge. Which verification approach fixes this?

contract testingpactschema registryintegrationfeedback loop
Pick one
Show the full answer Hide the answer

The deciding property

Two facts in the stem settle it. The breakage is found a day or two after the merge, in a shared environment — so the feedback loop, not the detection, is what is broken. And the breakages are in REST APIs as well as the event schema, so an event-only mechanism cannot cover the problem.

The fix has to move detection to the moment the provider's change is proposed, inside the provider's own pipeline, and it has to work for request/response APIs.

Why consumer-driven contracts

Each consumer declares what it actually uses from a provider: these fields, these types, this status code for this input. Those expectations are published, and the provider's pipeline verifies its new version against every consumer's recorded expectations before the change merges.

Three consequences make this the right shape here:

  • The feedback moves from two days to a few minutes, and lands on the person who made the change while they still have the context.
  • It tests what is used, not what exists. A provider can remove a field no consumer reads, which is the thing a schema-compatibility rule forbids and the thing teams most need to be able to do. Over twelve services and six teams, the ability to remove safely is what stops the API surface calcifying.
  • It scales with teams, not with paths. Twelve services have a quadratic number of interaction paths and a linear number of provider-consumer pairs. The contract approach grows with the smaller number.

What it costs

Be honest about the bill: a broker or shared store for the contracts, a verification step in every provider pipeline, and the discipline that a consumer's contract reflects real usage rather than an aspiration. Contracts also verify shape and semantics only as far as the consumer stated them — they will not catch a provider that still returns a correctly-shaped price with the wrong value. A thin end-to-end suite remains necessary for exactly that, which is why the answer is not "delete the E2E tests".

Why the other options fail

  • Expand the end-to-end suite. The most popular answer and the one that makes things worse. E2E in a shared environment is already the mechanism that is failing: it is slow, it is late, it requires every service deployed together, and its failures are ambiguous about which change caused them. Expanding it increases runtime and flakiness, which increases re-running, which removes the signal. It also cannot tell the provider's author anything before they merge, which is the actual requirement.
  • A central schema registry with compatibility checks. Genuinely good, and insufficient here. A registry enforces structural compatibility on messages — it will stop a required field being removed from the shared event schema, and that is real value. It does not cover the dozen REST APIs, and structural compatibility is a weaker guarantee than consumer expectations: a provider can satisfy the schema and still break a consumer by changing a status code, a pagination convention, or the meaning of a value. Run a registry and contracts; if forced to pick one for this situation, the registry covers a fraction of the stated problem.
  • Architecture review board for every cross-team API change. Adds days of latency to every change to catch something a machine can catch in seconds, and reviews by inspection miss the subtle cases precisely because they look compatible. It also makes the API surface harder to improve, so teams route around it.

When this is the wrong answer, and what would flip the decision

For two services owned by one team, contracts are pure overhead: the coordination cost they remove does not exist, and an integration test in one pipeline covers it. The practice earns its setup at roughly the point where the provider's author and the consumer's author are different people who do not sit together.

If this changes Choose Because
The integration is event-only, no REST Schema registry first It covers the whole surface for far less setup
Providers are external, outside your pipelines Consumer-side contract tests against a recorded stub, plus versioning You cannot verify in a pipeline you do not own
Two services, one team Neither; an integration test The coordination cost the contracts remove does not exist
The failures are wrong values, not wrong shapes A thin E2E suite on the critical paths Contracts verify the agreement, not the correctness of the data

What a strong answer adds

That the contracts must be verified against the provider's new version in the provider's pipeline, not merely written down. A contract that is published and never verified is documentation, and the failure mode of this pattern is a team that adopts the tooling, generates the files, and never wires the verification step into the build that blocks the merge.