The way back is also a deploy
How production teams undo a bad deploy, and the design-time contracts that decide whether rollback is an operation or a diagnosis.
Reconstructs rollback as it actually works in production: a forward deploy of an old artifact, underwritten by a compatibility contract fixed weeks before the incident. Six postmortems (Knight, CrowdStrike, Cloudflare, Reddit, GitHub, Azure AD) are grouped into four failure classes, and the doctrine of Amazon, Google, Meta, Netflix and Slack into five decisions with the conditions that flip them. A reader leaves able to say whether their own rollback is safe, automated and rehearsed, and what it would cost to make it so.
The most expensive deployment failure on record, Knight Capital, was a rollback that executed cleanly: it restored old code under new configuration and turned a one-server failure into an eight-server one.
What you get out of it
- There is no reverse gear: every published implementation rolls back by deploying the previous version forward, so pipeline health is rollback health.
- Rollback restores code, not context; flags, config, schema and serialized data must move with the artifact or the rollback amplifies the failure.
- Kubernetes never implemented the automatic rollback its docs promised for a decade; the ecosystem rebuilt it three times one layer up (Spinnaker, Argo Rollouts, Flagger).
- Automated rollback became defensible when detection did: Gandalf reports 92.4% precision and 100% recall gating Azure's data-plane rollouts.
- The rollback window closes at write speed: GitHub lost its failback after 40 minutes of diverged writes and paid 24 hours for the return trip.
Scope
Why this, now. CrowdStrike's 2024 RCA and Meta's 2026 health-check paper both landed remediations on the same point: the undo path is designed, not improvised, and most teams have never tested theirs.
What it does not cover. Restoring data from backups, live schema-migration mechanics, recalling software from uncontrolled devices, and incident process in general, each covered by a sibling guide.
Other field guides
Where resilience policy lives: Netflix, 2016 to 2026
Between 2016 and 2026 almost every Netflix library that decided something about the network was retired, while the libraries that decide something ab…
22 sources · 4 organisations · 4 postmortemsDeciding who gets told no
A field guide to request-level admission control, built from GitLab's public incident tracker and six years of its rate-limiting change record, the E…
26 sources · 13 organisations · 4 postmortemsAdding capacity under fire
Reconstructs, from postmortems at Slack, AWS, Datadog, Robinhood and Coinbase plus the Kubernetes project's own rejected pull requests, why the mecha…
28 sources · 18 organisations · 6 postmortems