advanced 3 min answer

Your public API has never made a breaking change by its own definition. In one release you reword an error message, start returning JSON object keys in a different order, and make a lookup 300 ms faster by caching it. Support tickets arrive from four integrators. What is the lesson, and what does an honest compatibility policy cost?

hyrums-lawimplicit-contractdeprecationapi-evolution
Show the full answer Hide the answer

What is being tested

Whether you can distinguish the contract you wrote from the contract your users actually have. With enough consumers, every observable behaviour of an interface becomes a dependency for somebody, whether or not you documented it. The observation is Hyrum Wright's and the name, Hyrum's Law, is Titus Winters's; it is set out in Software Engineering at Google (2020). The practical consequence is that "we only changed undocumented behaviour" is not a defence, it is a description of how the outage happened.

What each of the three changes broke

  • The reworded error message. An integrator was matching on the string, because the error code was too coarse to distinguish the case they needed. Their parser is awful and their need was real. The message changed, their branch stopped matching, and payments they used to retry now fail.
  • The key order. Someone was canonicalising the response body by serialising it and hashing it, as a cheap change-detection cache. JSON object key order is explicitly insignificant, and their hash changed for every record at once, so their nightly diff reported the whole dataset as modified.
  • The 300 ms improvement. A client was issuing a dependent write immediately after the read, relying on the old latency to hide a race in their own code. Faster made it visible. Performance is part of the observable surface, and improving it is a behaviour change.

What the honest policy costs

Accepting Hyrum's Law does not mean freezing. It means paying for change deliberately in four places:

  1. A machine-readable, stable error taxonomy so nobody has a reason to parse prose. A stable type or code field, documented, plus a human detail field explicitly marked as non-contractual. Then changing the prose is genuinely safe, because you gave the client something better.
  2. Signalled deprecation with real dates. The Deprecation response header (RFC 9745, March 2025) says a resource is being phased out; Sunset (RFC 8594, 2019) says when it stops answering. Both are machine-readable, so an integrator's monitoring can tell them before their customers do.
  3. Usage telemetry per consumer, per field, per version. You cannot run a brownout or an impact assessment for a change if you cannot name the four accounts that will feel it. This has to exist before the change, which is why it is almost always missing when it is wanted.
  4. A deliberate brownout rather than a cliff. Fail the deprecated behaviour for one minute an hour, with the error naming the deadline, weeks before removal. It converts a silent dependency into a support ticket at a time you chose.

The bill is real: a stable error contract constrains your internal refactoring, and telemetry at field granularity is work. What you buy is the ability to change anything at all without a 3 a.m. discovery.

When this is the wrong answer

For an internal API with three consumers in one repository, the fix is a grep and a conversation, and building deprecation infrastructure for it is waste. The investment scales with the number of consumers you cannot call. And some implicit dependencies should be broken on purpose: if a client is relying on a race window, closing it is correct, and the right move is a heads-up and a deadline rather than preserving a bug for them forever.