A developer-tools vendor of JetBrains' shape publishes a plugin API used by several thousand third-party plugins it does not control and cannot enumerate. An architect proposes consumer-driven contract tests. Which approach actually protects a release?
Show the full answer Hide the answer
The deciding property
Consumer-driven contract testing has one precondition that is easy to miss: the consumer set must be enumerable and willing. The mechanism is that each consumer publishes its expectations and the provider verifies against all of them. With several thousand consumers the vendor cannot list, the registry stops being the consumer set and becomes a self-selected sample.
That sample is biased in the worst possible direction. The consumers who have the capacity to publish contracts are the professional ones with their own test suites — the ones least likely to be broken silently. The breakage lands on the plugin written in a weekend in 2019 whose author has moved on.
Why replay works here
What the vendor cannot enumerate, it can observe. Instrument the API surface so every externally reachable symbol records which callers touched it, and build the release gate from that record: replay the observed call patterns against the candidate build and fail on a changed response shape or a changed outcome.
Weight the record by distinct consumers per symbol, not by call volume. Usage distributions here are brutally long-tailed: the top fifty calls are typically the overwhelming majority of traffic, while the removal risk sits in symbols called rarely by a handful of plugins each. A volume-weighted view declares those safe to delete.
Why the other options fail
- Twenty largest plugin authors. This is real work with real value for those twenty, and it protects the consumers you were already going to notice. It leaves the tail, which is where unenumerated consumers live, entirely uncovered. Right answer when the consumer set is twenty.
- Staging instance with sample plugins. The samples are written by the provider, so they encode the provider's beliefs about how the API is used. They pass exactly the usage you anticipated. This is the option that feels most like testing and tests the least.
- Freeze v1 and add v2. Versioning defers the problem rather than removing it: you will still change v1 for security and defect fixes, and a behaviour change breaks plugins with no schema change at all, which no amount of versioning catches. It also doubles the surface you maintain and verify.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| The consumers are internal teams on one release train | Consumer-driven contracts | They are enumerable and their build can be made to fail |
| The API is a managed partner programme with contracts | Both | Replay for the tail and published contracts for the partners you owe an SLA |
| You are adding rather than changing symbols | Neither | Additive change needs a compatibility lint and nothing more |
When this is the wrong answer
Replay cannot see call sequences nobody has made yet, and it cannot see consumers that depend on semantics rather than shape — the plugin relying on an undocumented ordering is invisible until it breaks. Treat replay as a floor, not a proof, and pair it with a deprecation path measured in releases rather than weeks, because for an unenumerable consumer set the only honest safety mechanism is time.