A release is burning 4% of checkout requests. The revert commit is merged at 14:06. At 14:08 your CI provider reports a degraded control plane and no new pipeline runs start. What happens next, and what should have been true?
Show the full answer Hide the answer
Second by second
14:06 the revert is merged and everyone relaxes, because the fix is "in". 14:08 the queue is not moving. 14:12 someone realises the organisation's time to recover is now bounded below by a third party's recovery, for which there is no commitment measured in minutes.
The root problem is in the first line: reversal was implemented as a forward build. A revert commit is a new tree that has never been compiled, never been tested and has no artifact, so undoing a release costs a full pipeline. The previous version's artifact already exists, already passed every stage and is sitting in the registry, and nothing in the recovery path refers to it.
Where it amplifies
The incident channel starts proposing direct changes to production. That is how a team that spent two quarters routing every change through the pipeline acquires an unmanaged change at the worst moment, applied by a tired person, with no record for the postmortem. Someone goes looking for the previous image digest and finds that deployments reference a mutable tag, so the question "what was running at 13:00" takes fifteen minutes to answer.
What the user sees
4% of checkouts failing for the duration of someone else's outage instead of the duration of a deployment. If the CI degradation lasts 90 minutes, the incident lasts 90 minutes, and the postmortem's timeline will show 14:06 as the moment the fix was ready.
What stops it
- Make reversal a selection, not a build. The deployment references an immutable digest, and rollback sets it to the previous digest. Seconds to a minute, no CI involvement, no compilation.
- Pin the last N known-good digests against registry retention, so a garbage-collection policy cannot delete your rollback target. Retention rules written for storage cost routinely do exactly that.
- Keep one deploy path with no CI dependency — a reconciler that already holds the manifests, or a single documented command with a stored credential — and exercise it quarterly, because an untested path is not a path.
- Use flags for behaviour so the first mitigation is a runtime change, not a release. A flag flip does not queue behind a build.
- Track time-to-deploy-with-CI-unavailable as a number. If nobody knows it, it is long.
What would have to be true for it to self-heal
Automated release verification that reverts on its own signals, where the revert action is the digest selection above. Automation layered on a rebuild inherits the same dependency and fails in the same minute, which is why the ordering matters: fix the reversal mechanism first, then automate it.
When this is not worth engineering for
A service that deploys weekly with a four-hour recovery objective: a CI outage fits inside the objective, and a tested out-of-band deploy path costs a credential to manage, a runbook to maintain and roughly a person-day a quarter to drill. Spend that where the recovery objective is minutes, where revenue moves per minute, or where the pipeline has already been the critical path in a real incident. Everywhere else, immutable digests and flags give most of the benefit for free.