concept

Reverse Compatibility Window

also called Revert Window, Rollback Precondition

The period during which every live consumer must still accept the previous version's output, because a rollback of one component is a forward-incompatible change for everything downstream of it.

rollbackcontractstolerant-readerversion-skewdead-letter

A billing service ships a version at 09:00 that requires a tax_basis field. The order service ships at 10:00 and starts sending it. At 10:40 the order release is causing checkout errors, so the on-call rolls orders back.

Orders stops sending the field and billing rejects every message. The rollback caused a worse failure than the one it was reversing, and because validation errors are 4xx, orders treats them as permanent and dead-letters the traffic instead of retrying.

Expand and contract is well understood as a way to roll a change forward safely. The same requirement applies in reverse and is routinely forgotten: a rollback is available only while every live consumer still accepts the previous producer's output. That period is the reverse compatibility window.

Why it matters

Rollback is the control most delivery practice depends on. Canary gates, bake times and progressive rollout all assume reverting is cheap and safe. If a consumer has already tightened its expectations, reverting the producer is a new outage and the safety story is fiction.

It fails silently, which is the dangerous part. Both services' error rates can look fine while every message is dead-lettered, and the user sees an order that succeeded and was never charged.

The window also decides whether an independent-deployment claim is true. Two services are independently deployable only if each can go back without the other's cooperation, which is a property of their contract, not of their pipelines.

Implementation patterns

  • Ship tolerant readers. A consumer accepts the field as optional with a defined default and makes it required only after the producer has emitted it for longer than the revert window.
  • State the window explicitly and derive it from policy: a 7-day revert policy needs 7 days of both-shape acceptance; a one-release revert policy needs one release cycle.
  • Never ship a newly required field in the same release train as the producer that supplies it. One train apart costs a week of feature latency and makes both sides individually revertible for free.
  • Enforce it in the pipeline: a consumer deploy that tightens a requirement refuses to proceed unless the producer's deployed version has been emitting the field for longer than the window.

Industry example

Kubernetes publishes this as a formal support matrix rather than a convention. Its version skew policy allows a kubelet to be up to three minor versions older than kube-apiserver from Kubernetes 1.28 (released in 2023) onward, having been two minor versions before 1.25. That is a declared reverse compatibility window: the control plane must keep serving nodes that have not moved, so an operator can upgrade the control plane and still roll a node pool back.

An upgrade whose reverse is unsupported turns a configuration mistake into a restore, and a documented skew window is what converts "we can go back" from a hope into a tested guarantee.

Failure scenarios

  • Producer rollback breaks the consumer, and the on-call rolls the consumer back too, putting two services on a pair nobody has run together.
  • A dead-letter replay spike: once the contract is repaired, tens of thousands of parked messages are replayed at once into a service sized for steady state.
  • Silent inconsistency: records created without their downstream effect, found in reconciliation hours later.

Trade-offs

The window costs feature latency and duplicated code paths. A field that could have been required on day one is optional for a release cycle, both branches are tested, and somebody has to remember to tighten it later, which frequently never happens.

What it buys is that rollback stays a minutes-long operation rather than a coordinated multi-team change under incident pressure. The exchange is a week of feature latency for the continued existence of your fastest recovery mechanism, which on a revenue path is not a close call.

When not to use it

In a single deployable this class does not exist. Producer and consumer change atomically and a rollback is one artefact going back, which is one of the real costs of splitting them in the first place.

If two services are owned by one team and always released together, the compatibility matrix is overhead and a shared release train is cheaper. For a contract with exactly one consumer you can redeploy in 90 seconds, a coordinated revert is reasonable as long as it is written down rather than improvised.

Interview question

Q: "You roll a service back and the incident gets worse. Walk me through the mechanism, and tell me what property you would have required before that release shipped."

What a strong answer covers: that reverting a producer is a forward-incompatible change for its consumers; the 4xx-and-dead-letter path and why neither error rate shows it; the precondition stated as a window derived from the revert policy; the tolerant-reader and release-train rules; and the note that a single deployable has none of this, so the window is a cost of the service split.

Quick check

Quiz: What single property has to hold for a rollback to be available? — Every live consumer still accepts the previous version's output, for at least as long as the revert policy's window.

Flashcard: Which signal reveals a contract break between two healthy-looking services? — The dead-letter rate. Validation rejections are 4xx, so the producer parks messages instead of retrying and neither error-rate dashboard moves.