Reliability & Operations 05 Oct 2026 28 min read

The way back is also a deploy

How production teams undo a bad deploy, and the design-time contracts that decide whether rollback is an operation or a diagnosis.

Reconstructs rollback as it actually works in production: a forward deploy of an old artifact, underwritten by a compatibility contract fixed weeks before the incident. Six postmortems (Knight, CrowdStrike, Cloudflare, Reddit, GitHub, Azure AD) are grouped into four failure classes, and the doctrine of Amazon, Google, Meta, Netflix and Slack into five decisions with the conditions that flip them. A reader leaves able to say whether their own rollback is safe, automated and rehearsed, and what it would cost to make it so.

The finding that surprised me

The most expensive deployment failure on record, Knight Capital, was a rollback that executed cleanly: it restored old code under new configuration and turned a one-server failure into an eight-server one.

What you get out of it

  • There is no reverse gear: every published implementation rolls back by deploying the previous version forward, so pipeline health is rollback health.
  • Rollback restores code, not context; flags, config, schema and serialized data must move with the artifact or the rollback amplifies the failure.
  • Kubernetes never implemented the automatic rollback its docs promised for a decade; the ecosystem rebuilt it three times one layer up (Spinnaker, Argo Rollouts, Flagger).
  • Automated rollback became defensible when detection did: Gandalf reports 92.4% precision and 100% recall gating Azure's data-plane rollouts.
  • The rollback window closes at write speed: GitHub lost its failback after 40 minutes of diverged writes and paid 24 hours for the return trip.

Scope

Why this, now. CrowdStrike's 2024 RCA and Meta's 2026 health-check paper both landed remediations on the same point: the undo path is designed, not improvised, and most teams have never tested theirs.

What it does not cover. Restoring data from backups, live schema-migration mechanics, recalling software from uncontrolled devices, and incident process in general, each covered by a sibling guide.

Open the field guide → Self-contained: it loads nothing at read time, follows your system theme, and prints cleanly.