An order service and a billing service deploy independently. Billing shipped a version at 09:00 that requires a tax_basis field on every message. Orders shipped at 10:00 and began emitting it. At 10:40 the orders release is causing checkout errors and the on-call rolls orders back. What happens next, and what property would have made that rollback safe?
Show the full answer Hide the answer
Second by second
10:40:30 — the old orders pods come up and stop emitting tax_basis.
10:40:45 — billing rejects the first message that lacks it. The rejection is a validation error, so billing returns a 4xx and does not retry; orders treats a 4xx as permanent and dead-letters the order.
10:41 — every order now fails at billing. The checkout error rate, which the rollback was supposed to fix, goes from a fraction of traffic to all of it. The rollback did not fail to help; it caused a worse failure than the one it was reversing.
10:43 — the on-call's instinct is to roll back further, so billing goes back to its pre-09:00 version too. That resolves the contract break and re-exposes whatever billing shipped at 09:00 to fix. Two services are now on versions nobody has run together today.
Where it amplifies
The dead-letter queue is the amplifier. Orders that failed are not lost, they are parked, so when the contract is repaired someone replays tens of thousands of messages at once into a billing service sized for steady state. The incident's second half is a self-inflicted load spike.
The second amplifier is organisational: two teams are now rolling back independently with no shared picture of which pairs of versions are compatible.
What the user sees
Not errors at first. Checkout appears to succeed, because orders accepted the request before billing rejected it. The user gets an order with no charge, and the discrepancy appears hours later as a missing payment. A contract break on an async hop presents as silent inconsistency, not as a 500.
What property would have made the rollback safe
A rollback is only available while every live consumer still accepts the previous producer's output. That is expand and contract applied in reverse, and it is a precondition, not a recovery step.
Concretely, billing should have shipped tax_basis as optional with a defined default, and made it
required only after orders had emitted it for longer than the rollback window. With daily deploys and a
policy of being able to revert to N-1, that window is one release cycle; with a 7-day revert policy it is
7 days of accepting both shapes.
The rule that prevents the whole class: never ship a newly required field in the same release train as the producer that supplies it. One train apart costs one week of latency on the feature and makes both sides individually revertible for free.
What stops it, and what would have to be true to self-heal
The mechanism is a tolerant reader on the consumer side plus a compatibility matrix the pipeline enforces: the deploy of a consumer that requires a field refuses to proceed unless the producer's currently deployed version has been emitting it for longer than the revert window.
For self-healing, billing would have to default the missing field rather than reject, and the retry would have to be transient rather than permanent. Neither is true here, which is why a human had to notice.
When this is the wrong answer
In a single deployable this class does not exist. The producer and consumer change atomically, and a rollback is one artefact going back. That is a genuine cost of splitting orders from billing, and it belongs in the decision: two services means every contract change is a two-phase change, forever. If the two services are owned by one team and always released together, the compatibility matrix is overhead and a shared release train is the cheaper control.
The signal to watch is not either service's error rate, which looked fine on both dashboards. It is the dead-letter rate, which is the only place a contract break between two healthy services becomes visible.