Undoing a bad deploy  / field guide
Practitioner field guide · 2026-10-05

The way back is also a deploy

How production teams undo a bad change, reconstructed from six published incidents and the deployment doctrine of Amazon, Google, Meta, Netflix and Slack. What rollback actually is mechanically, the four ways it fails, and the decisions, made weeks before the incident, that keep the way back open.

30 primary sources 11 production systems 6 incidents Evidence through October 2026 Read: 22 min
01

The territory

The problem, who has solved it in production, and the one fact about rollback that the famous incidents keep re-teaching.

$460M
Lost in 45 minutes after a rollback made the failure eight servers wide instead of one
8.5M
Windows hosts crashed by one content update whose revert could not reach machines that no longer booted
314min
Reddit's Pi-Day outage; hours of it were spent deciding whether the untested rollback was riskier than staying down
~70%
Of outages trace to a change in a live system, which is why the undo path is the one that matters

Strip the vocabulary away and the problem is this: a change you made is hurting production, and you must return the system to its last known-good behaviour faster than the damage compounds. The return trip is itself a change. It runs through the same machinery that just failed you, it is executed at the worst possible moment, under the least information and the most pressure, and it is almost always the least rehearsed operation the team owns.

The organisations that have written seriously about this problem solved it from both ends. Amazon's Builders' Library states the doctrine most plainly: every revision must be backwards compatible, meaning it "can be rolled back without errors or disruption", and that property is verified in the pipeline, not assumed, with an upgrade-then-rollback test run against part of the fleet before full deployment. Google's SRE book, Meta's health-check paper, Microsoft's Gandalf paper and Slack's ReleaseBot account all describe the same structure from different angles: expose the change progressively, watch an independent health signal, and make the return trip automatic because, in Google's words, humans in repetitive loops bring fatigue and inattention.

The fact the incidents keep re-teaching is the one this guide leads with. The most expensive deployment failure in the public record, Knight Capital, was not a failed rollback. The rollback executed cleanly. According to the SEC's order, engineers uninstalled the new code from the seven servers where it had been deployed correctly, and because the new code had repurposed an old feature flag, the rollback re-activated eight-year-old dead code on all eight servers instead of one. Rollback restores your code. It does not restore the context your code runs in: the flags, the configuration, the data, the traffic. Every failure class in section 4 is a variation on that asymmetry.

Scope, stated up front: this guide covers getting a bad application deploy, config push or content update out of production, and the schema and state contracts that make that possible. It deliberately does not cover restoring data from backups (see the sibling guide on proving you can restore), the mechanics of live schema migration (covered in changing the schema under load), or recalling software from devices you do not control (updating software you cannot recall). Incident process in general is in learning from production failure.

Figure 1 · Four ways back, and what each costs

seconds

seconds to minutes

minutes, via the pipeline

hours, often manual

Bad change
live in production

Flip the flag or
kill switch

Shift traffic back
to the stable fleet

Redeploy the
previous artifact

Repair or restore
diverged state

Blast radius
stops growing

seconds

seconds to minutes

minutes, via the pipeline

hours, often manual

Bad change
live in production

Flip the flag or
kill switch

Shift traffic back
to the stable fleet

Redeploy the
previous artifact

Repair or restore
diverged state

Blast radius
stops growing

Time-to-undo grows left to right, and everything to the right of a traffic shift carries state risk. The fastest undo is the one that never touches the artifact. Reconstructed from AWS, Slack and Savor et al.
Diagram source
02

How it is actually built

The shape that recurs across Amazon, Meta, Azure, Netflix and Slack, and the mechanical fact about rollback that explains most of the failures.

Start with the mechanism, because it is widely misunderstood. There is no reverse gear. In every system with a published implementation, rolling back means deploying the previous version forward. The Kubernetes deployment controller's rollback.go literally "compares the templates of the provided deployment and replica set and updates the deployment with the replica set template in case they are different": the old pod template is copied onto the Deployment and an ordinary rollout proceeds. Slack's account of ReleaseBot says the same thing from the operations side; per the getdx interview with Sean McIlroy, rollback is Slack's fastest deploy, taking a minute or two, precisely because it is a deploy with the build and test steps already paid for.

That one fact carries three consequences. First, everything that makes your deploys slow makes your rollbacks slow; a 40-minute pipeline is a 40-minute undo. Second, rollback inherits every dependency of the forward path, so when the forward machinery is what broke, the way back is blocked too: CrowdStrike's revert shipped 78 minutes after the bad content, but hosts that had already loaded it were boot-looping and could not fetch anything, and Azure AD's automated rollback in September 2020 failed because the same defect that mis-targeted the deployment had corrupted the metadata the rollback needed. Third, whether the old version still works is not a property of the rollback tool at all. It is a property of the compatibility contract between versions, which was fixed weeks earlier, when the change was designed. Amazon's doctrine makes this explicit: rollback safety is decided at design time and verified in pre-production, because by the time you need it, it is too late to add.

Figure 2 · The reference shape

Independent health signal

Progressive exposure

go

no-go

underwrites

Build and
preprod suites

One-box or canary
2-10% of traffic

Waves or rings,
each with bake time

Alarms, z-scores,
canary statistics,
templated health checks

Go / no-go
gate

Promote the
next wave

Rollback executor:
deploy the previous version

Fleet back on
last known good

Compatibility contract:
two-phase deploys,
expand/contract schema

Independent health signal

Progressive exposure

go

no-go

underwrites

Build and
preprod suites

One-box or canary
2-10% of traffic

Waves or rings,
each with bake time

Alarms, z-scores,
canary statistics,
templated health checks

Go / no-go
gate

Promote the
next wave

Rollback executor:
deploy the previous version

Fleet back on
last known good

Compatibility contract:
two-phase deploys,
expand/contract schema

The common structure across AWS pipelines, Meta's SHC, Azure's Gandalf, Netflix Kayenta and Slack ReleaseBot. The dashed edge is the part most teams skip: the executor only works if the compatibility contract held.
Diagram source

Progressive exposure

Amazon's one-box stage serves at most ten percent of a zone's requests before wave deployment begins, per Clare Liguori's account. Slack sends builds to a dogfood tier, then a canary taking about 2% of production traffic. Azure deploys in rings with bake time between them. The point is not caution for its own sake; it is that the undo cost in figure 1 is proportional to exposure at detection time.

Runs this way at: AWS, Slack, Azure

The independent health signal

Every implementation separates the thing that deploys from the thing that judges. Slack's ReleaseBot watches z-scores on error and latency series. Kayenta compares canary and baseline statistically and fails the canary on significant degradation. Meta's Service Health Checker runs templated metric checks wired into every phased rollout. Gandalf goes furthest: an ensemble that attributes fault signals to one of many concurrent rollouts before deciding no-go.

Sources: InfoQ on Slack, Netflix, Meta, Microsoft

The gate and the executor

Azure's deployment engine subscribes to Gandalf's go/no-go decisions and stops a rollout on no-go. AWS pipelines roll back the one-box automatically without promoting further. The executor itself is boring by design: deploy the previous artifact, shift traffic back, or both. Boring is the feature. The executor runs on the worst day, so every clever branch in it is a liability.

Sources: Gandalf, NSDI 2020, AWS

The compatibility contract

The load-bearing component has no runtime presence at all. Amazon requires serialized data, APIs and stored state to be readable by version N when written by version N+1, enforced through two-phase deployments (figure 3) and verified by deploying to half a fleet, rolling back, and checking both directions. Without this contract the executor is a loaded gun: the artifact reverts, the data does not.

Source: AWS Builders' Library, 2019

Figure 3 · The two-phase deployment keeps every rollback one safe step

rollback safe

rollback safe

rollback safe

Version N · writes old format, reads old only

Version N+1, prepare · writes old format, reads old and new

Version N+2, activate · writes new format, reads old and new

Version N+3, clean up · drops the old read path

rollback safe

rollback safe

rollback safe

Version N · writes old format, reads old only

Version N+1, prepare · writes old format, reads old and new

Version N+2, activate · writes new format, reads old and new

Version N+3, clean up · drops the old read path

Amazon's rule for serialized data, per the Builders' Library: the reader learns the new format one release before any writer uses it, so no rollback ever lands on data it cannot parse. Expand/contract is the same protocol applied to schemas.
Diagram source

Where implementations diverge is instructive. Who decides varies from fully automatic (AWS pipelines, Gandalf, ReleaseBot after 2024) to advisory (Kayenta can route to a human approval path) to fully manual (Reddit's operators, on the day). What reverts first varies: Facebook and OANDA, per Savor et al., reach for config kill switches before artifact rollbacks, while Cloudflare's 2019 recovery was a global kill switch on the WAF ruleset, not a redeploy. And where the capability lives varies most of all, which is section 3's second decision.

03

The decisions that matter

Four forks, each with the condition that flips the answer. The pattern underneath all four: rollback is a capability you buy at design time, not a button you press at incident time.

When a deploy goes bad, is the first move rollback or fix forward?

Chosen
  • Rollback as the reflex. Google's SRE book pairs progressive rollout with "rolling back changes safely when problems arise". Slack's ReleaseBot rolls back on anomaly, and asks questions afterwards.
Rejected
  • Diagnose-then-patch under incident pressure. A fix forward needs a correct diagnosis, a build and a pipeline run at the exact moment you have the least of all three. Knight tried to understand the bug with the market open.
Flips when
  • State has moved under the new version (a migration ran, new-format data is written) and old code no longer runs against it. Then rollback is the dangerous move, and a minimal revert through an emergency pipeline is safer.

Does a machine or a human pull the trigger?

Chosen
  • Automatic, metric-gated rollback at AWS, Meta, Netflix and Slack. The enabling number is detection quality: Gandalf reports 92.4% precision and 100% recall on data-plane rollouts, which is what makes handing the trigger to software defensible.
Rejected
  • Per-deploy human judgement. Slack ran deploy commanders in two-hour shifts and concluded the vigilance does not scale; Google's stated reason is fatigue and inattention in repetitive tasks.
Flips when
  • Your false-positive rate would make the bot cry wolf: new services without baselines, bursty metrics, or rollouts too rare to tune against. Meta's paper is candid that noise, drift and alert fatigue were the actual engineering cost.

Does rollback live in the platform core or a layer above?

Chosen
  • Kubernetes chose the layer above, three separate times by the ecosystem: Spinnaker, Argo Rollouts and Flagger all exist partly because, in Argo's words, a Deployment "can halt the progression, but [is] unable to automatically abort and rollback the update".
Rejected
  • Auto-rollback in core. Requested in issue #23211 in 2016, left in the SIG Apps backlog, never planned. The server-side rollback API (rollbackTo) was deprecated in favour of a client-side kubectl rollout undo.
Flips when
  • Your source of truth is git. Under GitOps the platform's rollback machinery is redundant by construction: revert the commit and let the controller converge. The orchestrator only needs to make forward deploys safe to repeat.

Do database changes ship with down migrations?

Chosen
  • Forward-only migrations plus expand/contract, so the application rollback stays safe while the schema never reverses. Ariga's position piece argues the down file is a myth of symmetry; AWS's two-phase rule is the same protocol applied to serialized data.
Rejected
  • Symmetric down scripts in production. A failed up-migration leaves the database in a state the down file does not expect, and the inverse of an additive change is destructive: dropping the column deletes the column's data. Octopus adds the decay argument: scripts nobody ever runs rot.
Flips when
  • Outside production. In dev and review environments down migrations are genuinely useful, which is Ariga's own carve-out, and tooling-managed down with checks is their product answer. The debate is about untested SQL run during an incident, not about the files existing.

Figure 4 · What to undo, decided in order

yes

no

no

yes

yes

no

A deploy is making
production worse

Is the behaviour behind a
flag or kill switch?

Flip it off.
Artifact untouched,
undo in seconds

Has state moved under it?
Migration ran, or new-format
data was written

Redeploy the previous
artifact through the
normal pipeline

Does the old version still
run against current state?

Roll back the artifact now,
contract the schema later

Fix forward: revert the
commit, smallest possible diff,
emergency pipeline

yes

no

no

yes

yes

no

A deploy is making
production worse

Is the behaviour behind a
flag or kill switch?

Flip it off.
Artifact untouched,
undo in seconds

Has state moved under it?
Migration ran, or new-format
data was written

Redeploy the previous
artifact through the
normal pipeline

Does the old version still
run against current state?

Roll back the artifact now,
contract the schema later

Fix forward: revert the
commit, smallest possible diff,
emergency pipeline

The order matters: the flag question is first because it is the only undo measured in seconds, and the state question is what separates a safe artifact rollback from a Knight-style amplification. Terminal nodes are actions, not advice.
Diagram source
DecisionChosenRejectedBecauseEvidence
First moveRoll back, reflexivelyFix forward under pressureThe incident is the worst time to diagnoseSRE book, Slack
TriggerAutomated, metric-gatedHuman per-deploy vigilanceFatigue; detection precision now supports itGandalf, Meta SHC
Where it livesLayer above the orchestratorAuto-rollback in Kubernetes coreCore keeps forward-only primitives; policy varies too much#23211, Argo
SchemaForward-only + expand/contractDown migrations in productionAsymmetry: reversing an additive change destroys dataAtlas, AWS
Exposure unitOne-box / canary firstFleet-wide pushUndo cost is proportional to exposure at detectionLiguori, Slack
The hinge

Every decision above reduces to one question: how old is the newest thing the previous version cannot understand? If the answer is "nothing, by contract" you may roll back blind, automatically, in minutes. The moment flags, config, schema or serialized data escape that contract, the undo stops being an operation and becomes a diagnosis.

04

What broke in production

Six published incidents, four failure classes. None of them is a rollback tool malfunctioning; every one is an assumption about the way back that nobody had tested.

Grouping the public record, four classes cover it. Class 1: rollback restores code, not context, the Knight asymmetry. Class 2: the undo nobody rehearsed, where the rollback was possible but the team could not price its risk, so the decision itself became the outage. Class 3: the undo needs the system to be alive, where the rollback depends on exactly the machinery the failure took out. Class 4: state closes the window, where every minute of forward operation makes the return trip more expensive, until it is no longer a rollback but a reconciliation project.

Figure 5 · Knight Capital: the rollback that armed the trap

MarketServer 8, old PowerPeg codeServers 1-7, newRLP codeOperationsMarketServer 8, old PowerPeg codeServers 1-7, newRLP codeOperationsDeploy copied the new code to only 7 of 8 servers. The activationflag was reused from Power Peg.Under fire, the newcode is suspected45 minutes, about$460M pre-tax lossenable flag (runs new RLP code)enable flag (activates dormant Power Peg)uncontrolled child ordersrollback, uninstall new RLP codePower Peg now live on all eight servers
MarketServer 8, old PowerPeg codeServers 1-7, newRLP codeOperationsMarketServer 8, old PowerPeg codeServers 1-7, newRLP codeOperationsDeploy copied the new code to only 7 of 8 servers. The activationflag was reused from Power Peg.Under fire, the newcode is suspected45 minutes, about$460M pre-tax lossenable flag (runs new RLP code)enable flag (activates dormant Power Peg)uncontrolled child ordersrollback, uninstall new RLP codePower Peg now live on all eight servers
The sequence per the SEC's 2013 order. The rollback was executed correctly and turned a one-server failure into an eight-server one, because the flag that activated the new code also activated the old dead code. Source: SEC release 34-70694.
Diagram source
Postmortem

Class 1 · Knight Capital, 2012

AssumptionRemoving the new code returns the system to yesterday's behaviour.
What happenedThe new router code reused a dormant feature's flag. Deploy reached 7 of 8 servers with no second-person check. The rollback removed the new code while the flag stayed on, so the dormant code ran everywhere, per the SEC's order.
Blast radiusAbout $460M pre-tax loss in roughly 45 minutes; the firm needed rescue financing within the week.
FixFor the industry: SEC Market Access enforcement. The firm did not survive independently.
Design ruleCode, flags and config are one deployable unit; roll them back together or not at all. Never repurpose a flag: a flag's meaning is part of the fleet's state.
Postmortem

Class 2 · Reddit Pi-Day, 2023

AssumptionWe will never need to roll back a Kubernetes control-plane upgrade, so the restore path can stay untested.
What happenedUpgrading 1.23 to 1.24 silently deleted the legacy "master" node label. Calico route reflectors were selected on that label by years-old, unversioned configuration, and the cluster network degenerated.
Blast radius314 minutes of global outage. A large share was decision latency: per the postmortem, it took hours to decide that a rollback, "a high-risk action on its own", beat staying down.
FixRestore from backup succeeded; remediations centred on making restores and upgrade rollbacks rehearsed, versioned and boring.
Design ruleAn untested rollback is priced at incident time, and the pricing is the outage. Rehearsal buys you the hours the decision would otherwise cost.
Postmortem

Class 3 · CrowdStrike Channel File 291, 2024

AssumptionA fast revert bounds the damage of a bad content push.
What happenedA 21-field template was fed 20 values; the out-of-bounds read crashed the Windows sensor in the boot path. Content went to the whole fleet at once. The RCA's timeline: pushed 04:09 UTC, reverted 05:27 UTC, but a machine that cannot boot cannot download a revert.
Blast radiusAbout 8.5M Windows hosts (Microsoft's estimate); Parametrix modelled $5.4B direct losses in the Fortune 500 alone. Recovery was per-machine, partly physical.
FixThe RCA's remediations are rollout remediations: staged deployment rings for content, customer control over timing, validation layers. The rollback lesson became a rollout lesson.
Design ruleRollback is a message, and a message needs a live receiver. For anything that can kill its own transport (boot paths, agents, network config), the undo must exist out of band, or exposure must be staged so the dead population is small.
Postmortem

Class 3 · Azure AD, September 2020

AssumptionThe safe-deployment system that stages our rollouts will also execute our rollbacks.
What happenedPer Microsoft's RCA as reported at the time, a latent defect in the Safe Deployment Process misread deployment metadata, so an update aimed at the test ring hit all rings concurrently; the same corruption then made the automated rollback fail, forcing a manual, ring-by-ring rollback.
Blast radiusRoughly five hours of authentication failures across Microsoft 365, Teams and the Azure portal, Americas worst affected.
FixMicrosoft's stated remediations: fix the SDP defect and expand resilience of the rollback path itself.
Design ruleThe guardrail is also software. A deployment-safety system that shares state with the deployment it protects is a single failure domain; test the rollback machinery with the same hostility as the service.
Postmortem

Class 4 · GitHub, October 2018

AssumptionIf failover misfires, we can fail back.
What happenedA 43-second network partition promoted West Coast MySQL primaries. By detection, the West had ingested about 40 minutes of writes while the East held seconds of unreplicated writes of its own; per the analysis, neither side contained the other, so failing back would have destroyed data.
Blast radius24 hours 11 minutes of degraded service, the time to restore, replay and reconcile rather than the time to fail over.
FixGitHub chose integrity over recovery time as explicit policy, then invested in cross-region replication and orchestration guardrails.
Design ruleThe rollback window is a function of write divergence, and it closes at write speed. If you intend to return, you must either fence writes immediately or accept that after N minutes the way back is gone.
Postmortem

Class 1 · Cloudflare WAF, July 2019

AssumptionRules must reach the whole edge in seconds, so the rule channel may skip staged rollout.
What happenedOne WAF rule with catastrophic regex backtracking shipped globally through the fast channel, and a CPU limit that would have contained it had been removed in an earlier refactor. CPUs saturated worldwide.
Blast radius27 minutes, 13:42 to 14:09 UTC, with traffic dropping about 82% at the trough.
FixRecovery was a global kill switch on the managed ruleset, not a redeploy; afterwards, staged rollout and simulation for rule changes, and the CPU guard restored.
Design ruleAny channel exempted from progressive rollout for speed must carry its own kill switch of equal speed. The exemption and the undo are one decision.

Figure 6 · Azure AD: when the rollback machinery shares the defect

All production ringsTest ringSafe DeploymentProcessReleaseAll production ringsTest ringSafe DeploymentProcessReleaseabout 5 hours ofglobal authenticationfailuresupdate targeted at the test ringonlydefect misreads metadata, all rings targeted concurrentlybackend services crashon startuptrigger automated rollbackfails, the same defect corrupteddeployment metadatamanual rollback, ring by ring
All production ringsTest ringSafe DeploymentProcessReleaseAll production ringsTest ringSafe DeploymentProcessReleaseabout 5 hours ofglobal authenticationfailuresupdate targeted at the test ringonlydefect misreads metadata, all rings targeted concurrentlybackend services crashon startuptrigger automated rollbackfails, the same defect corrupteddeployment metadatamanual rollback, ring by ring
Reconstructed from Microsoft's RCA as quoted by contemporaneous reporting. One latent defect produced both the mis-targeted deployment and the failed automated rollback, which is what made a ring-staged system fail globally. Sources: BleepingComputer, The Register.
Diagram source
05

Numbers you can plan against

What undo speed, exposure and failure actually measure out to in the published record.

MetricValueAtContextAs ofSource
Outages caused by live-system changes~70%GoogleStated in the SRE book's change-management discussion; internal, undated measurement2016SRE book
Rollback execution time1-2 minSlackReleaseBot's rollback, described as the fastest deploy they run2024getdx interview
Canary exposure before promotion~2%SlackShare of production traffic on the canary tier2020Deploys at Slack
One-box exposure cap10%AmazonMaximum share of requests served by the one-box stage2021Liguori
No-go detection precision92.4% / 100% recallAzureGandalf, data-plane rollouts over 18 months in production2020NSDI 2020
Bad-deploy loss rate$460M / 45 minKnightPre-tax loss between flag-on and full stop, rollback included2012SEC order
Push-to-revert latency78 minCrowdStrike04:09 UTC push to 05:27 UTC revert; derived from the RCA timeline2024RCA
Hosts crashed before revert landed~8.5MWindows fleetMicrosoft's estimate, under 1% of Windows machines2024Cybersecurity Dive
Modelled Fortune 500 direct losses$5.4BParametrixInsurer's model, average $44M per affected company; claimed, not audited2024Cybersecurity Dive
Decision latency dominating an outage314 min totalRedditHours of it spent deciding whether the untested rollback was safe2023postmortem
Divergence that closed the failback window~40 min of writesGitHubUnreplicated writes on both sides after a 43-second partition; 24h11m to full recovery2018GitHub
Deploy frequency under this discipline~12/daySlackScheduled webapp deploys, each a potential rollback rehearsal2020Deploys at Slack
Read these carefully

Measured: the SEC, CrowdStrike, GitHub, Reddit and Cloudflare figures come from primary incident documents. Self-reported: Gandalf's precision and recall, Slack's rollback time and Amazon's exposure caps are the operators describing their own systems, with no independent measurement published. Modelled: Parametrix's $5.4B is an insurer's estimate and should be quoted as such. Undated: Google's 70% has no published methodology or date; treat it as directional, not as a planning constant.

06

The evidence wall

Every source behind this page, graded. Filter by kind. Links were retrieved 2026-10-05; this build's container could re-verify GitHub-hosted links directly, others via server-side retrieval only.

Postmortem U.S. SEC / Knight Capital2013-10

In the Matter of Knight Capital Americas LLC (34-70694)

The regulator's reconstruction: code deployed to 7 of 8 servers with no second reviewer, a repurposed flag, and a rollback that activated dead code fleet-wide.

Carry forwardRollback restores code, not context. Flags and config are state.
sec.gov
Postmortem CrowdStrike2024-08

Channel File 291 external root cause analysis

21 fields defined, 20 supplied; the revert shipped in 78 minutes and could not reach hosts that no longer booted. Remediations are all rollout-shaped: rings, validation, customer control.

Carry forwardAn undo is a message; it needs a live receiver.
crowdstrike.com
Postmortem Cloudflare2019-07

Details of the Cloudflare outage on July 2, 2019

A WAF rule skipped staged rollout by design and saturated every CPU on the edge. Recovery came from a global kill switch, not a redeploy.

Carry forwardA fast-path channel earns a kill switch of equal speed.
blog.cloudflare.com
Postmortem Reddit2023-03

You Broke Reddit: the Pi-Day outage

A Kubernetes upgrade deleted the node label years-old unversioned config depended on. The restore worked; pricing an untested restore mid-incident cost hours.

Carry forwardRehearse the undo so the decision is cheap when it matters.
reddit.com/r/RedditEng
Postmortem GitHub2018-10

October 21 post-incident analysis

43 seconds of partition, 40 minutes of diverged writes, and the failback was gone. GitHub chose integrity over recovery time and paid 24 hours for it, explicitly.

Carry forwardThe rollback window closes at write speed; fence writes or forfeit the return.
github.blog
Postmortem Microsoft (via press)2020-10

Azure AD outage RCA: the SDP defect

One latent defect mis-targeted all rings at once and corrupted the metadata the automated rollback needed. Cited through BleepingComputer and The Register, which quote the RCA; Microsoft's own status-history page is not stably linkable.

Carry forwardThe deployment guardrail is software too; test it like the service.
bleepingcomputer.com
Design doc AWS / Sandeep Pokkunuri2019

Ensuring rollback safety during deployments

Amazon's doctrine: backwards compatibility at every revision, the two-phase deployment for serialized data, and the upgrade-rollback test as a pipeline gate.

Carry forwardRollback safety is designed and tested before deploy, never improvised after.
aws.amazon.com
Design doc AWS / Clare Liguoriundated

Automating safe, hands-off deployments

The pipeline anatomy: one-box capped at 10% of requests, wave deployment, bake time, alarm-gated promotion, automatic rollback with no human in the loop.

Carry forwardBake time is the budget your detection needs to catch slow-burn regressions.
builder.aws.com
Source Kubernetes2016-03

Issue #23211: automatic rollback for failed deployments

The feature request, probe-driven auto-undo with a failure budget, parked in the SIG Apps backlog and closed without implementation.

Carry forwardThe orchestrator keeps forward-only primitives; rollback policy lives above.
github.com
Source Kubernetes2026-04

Website issue #55312: the promise that never shipped

The docs said the controller would roll back automatically "in the future, once automatic rollback will be implemented". A decade later the sentence was removed rather than the feature added (PR #55402).

Carry forwardAudit what your platform's docs promise about undo; some promises are aspirational.
github.com
Source Kubernetes2020-08

PR #93807, approved then closed unmerged

The cleanup removing dead server-side rollback logic was approved in two days, went stale waiting on a rebase, and the bot closed it. The deprecated code path outlived the attempt to delete it.

Carry forwardRejected and rotten PRs record what a project will not prioritise; read them before assuming.
github.com
Source Kubernetescurrent

pkg/controller/deployment/rollback.go

The implementation: rollback copies the old ReplicaSet's template onto the Deployment and rolls forward; the old API survives only as a deprecated annotation that is cleared after use.

Carry forwardThere is no reverse gear; a rollback is a forward deploy of an old artifact.
github.com
Source Argo projectcurrent

Argo Rollouts README

States the gap it fills in one line: a Deployment "can halt the progression, but unable to automatically abort and rollback the update". Features: automated rollbacks and promotions, metric analysis.

Carry forwardIf the platform will not undo, the layer above will be built, three times over.
github.com
Source Flux CDcurrent

Flagger README

Progressive delivery operator: shifts traffic gradually while measuring metrics and running conformance tests, reverting on threshold breach.

Carry forwardTraffic shift is an undo mechanism in its own right, cheaper than redeploying.
github.com
Case study Google2016

SRE book, chapter 1: change management

The 70% figure and the prescription: progressive rollouts, fast accurate detection, safe rollback, humans out of the repetitive loop.

Carry forwardIf ~70% of outages are changes, undo speed is availability work, not tooling polish.
sre.google
Case study Parametrix via Cybersecurity Dive2024-07

CrowdStrike losses in the Fortune 500

The insurer's model: $5.4B direct losses, $44M average per affected company, 10-20% insured. A modelled estimate, labelled as such here.

Carry forwardThe cost of a dead undo path is measured in other companies' balance sheets.
cybersecuritydive.com
Eng blog Slack2020

Deploys at Slack

Twelve scheduled deploys a day, a rotating deploy commander, dogfood then a 2% canary, rollback and hotfix as the two standing remedies.

Carry forwardDeploy frequency is rollback rehearsal; rare deploys mean untested undos.
slack.engineering
Eng blog Slack / Sean McIlroy2024

The scary thing about automating deploys

ReleaseBot replaced the human deploy commander: z-score anomaly detection, automatic pause and rollback. The getdx interview adds the number: rollback is their fastest deploy, one to two minutes.

Carry forwardThe emotional blocker to automated rollback is trust in detection; earn it with precision data.
slack.engineering
Eng blog Netflix2018-04

Automated canary analysis with Kayenta

Statistical comparison of canary and baseline metrics; on significant degradation the canary is aborted and traffic routes back to stable. Built with Google, open-sourced into Spinnaker.

Carry forwardJudge the canary against a fresh baseline replica, not against all of production.
netflixtechblog.com
Eng blog Ariga (Atlas)2024-04

The myth of down migrations

The argument that symmetric down files cannot deliver what they promise: failed migrations leave unknown states, and reversing additive changes destroys data.

Carry forwardPlan schema undo as expand/contract phases, not as inverse SQL.
atlasgo.io
Eng blog Pragmatic Engineer2023-04

Learnings from the Reddit outage

Independent close reading of the Pi-Day postmortem, preserving the line this guide leans on: hours to decide that a rollback, a high-risk action on its own, was the best course.

Carry forwardDecision latency is a measurable, improvable part of time to recover.
pragmaticengineer.com
Vendor Octopus Deployundated

Pitfalls with SQL rollbacks and automated database deployments

The decay argument from a deployment vendor: rollback scripts that are never run are wasted work, and the motivation to write them falls with every successful deploy.

Carry forwardAn undo artifact that is never exercised converges on being wrong.
octopus.com
Paper Savor et al., ICSE2016-05

Continuous deployment at Facebook and OANDA

The field study: remedies for a bad deploy in practice are rollback, rapid hotfix, or a config change disabling the feature, in that menu, across two very different firms.

Carry forwardKeep all three remedies stocked; each has a failure mode the others cover.
dl.acm.org
Paper Li et al., NSDI2020-02

Gandalf: safe deployment in large-scale cloud infrastructure

Azure's go/no-go analytics: ensemble fault attribution across concurrent rollouts, 92.4% precision and 100% recall on the data plane over 18 months, wired directly into the deployment engine.

Carry forwardAttribution across concurrent changes is the hard part of auto-rollback at scale.
usenix.org
Paper Meta, ISSRE industry track2026-08

Making deployments safe at Meta: health checks for continuous change-safety

The Service Health Checker: templated metric checks composed by service owners, integrated with tiered and phased rollouts to trigger automatic rollback on regressions, and a frank account of noise, drift and alert fatigue at fleet scale.

Carry forwardDefault health checks beat bespoke ones; the failure mode is silence, not noise.
arxiv.org
Paper Klein, Betser, Monroe, ;login:2014-10

Making "push on green" a reality

Google's early public account of mechanised rollouts: if the tests are good and the build is good, push, with rollback as a first-class, rehearsed step rather than an escape hatch.

Carry forwardAutomated promotion and automated rollback are one system; adopting half of it is the risky configuration.
usenix.org
Talk Clare Liguori, InfoQ podcast~2021

Automating safe and hands-off deployments at AWS

The spoken companion to the Builders' Library article: one-box stages, alarm-gated waves, automatic rollback as the default behaviour of every pipeline.

Carry forwardNo human watches an Amazon deploy by default; the pipeline owns the undo.
infoq.com
Talk Peter Ramensky, AWS re:Invent2019-12

Amazon's approach to high-availability deployment (DOP404)

The conference treatment of fractional deployments, traffic shifting, bake time and rollback. Cited from the published deck; this build's environment could not retrieve the video for timestamped claims.

Carry forwardFractional exposure plus automatic rollback is presented as one indivisible practice.
d1.awsstatic.com
07

Build a miniature, then productionise it

Seven rungs from a toy blue-green flip to breaking your own rollback machinery on purpose. The crossing from toy to real is rung four.

Two versions, one switch

Run v1 and v2 of a trivial service behind nginx or Caddy; flip traffic between them with a config reload. Script the flip both ways.

Done when: either direction completes in under 5 seconds with zero dropped requests under light load.  Teaches: traffic shift as the cheapest undo.

Make the machine pull the trigger

Ship a deliberately broken v2 (10% error rate). A watchdog polls the error ratio and flips traffic back without you.

Done when: injected failure to full recovery is under 60 seconds, unattended, ten runs in a row.  Teaches: the gate-and-executor loop, and your first false-positive tuning.

Let state poison the rollback

Have v2 write a new field into a queue message or cache entry that v1 cannot parse. Roll back and watch v1 crash on v2's leftovers.

Done when: you can reproduce the post-rollback failure on demand and explain exactly which artifact carried the poison.  Teaches: the Knight asymmetry at toy scale.

Fix it with a two-phase deploy

Re-ship the same change as two releases: first teach the reader to accept both formats, then start writing the new one. Roll back from each phase.

Done when: rollback from every intermediate state leaves the system healthy, demonstrated by test.  Teaches: the compatibility contract that underwrites every real rollback.

Expand/contract a real schema

Add a column the new code uses while the old code still runs, then contract only after the rollback window closes. No down migration anywhere.

Done when: v1 runs correctly against the fully migrated schema, proven in CI.  Teaches: why forward-only migration is a rollback feature, not a limitation.

Put the rollback test in the pipeline

Amazon-style gate: CI deploys the new build to half your toy fleet, rolls it back, and asserts health in both directions before anything promotes.

Done when: a commit that breaks backward compatibility of stored data is rejected by the pipeline, not by a human.  Teaches: rollback safety as a testable property.

Break the way back on purpose

Game-day your own machinery: corrupt the artifact store, kill the config service mid-rollback, make the previous image unpullable. Measure time to last known good.

Done when: you have a measured number for "time to previous version with the happy path blocked" and a written fallback.  Teaches: the Azure AD lesson; the guardrail is also software.

08

Keep hunting

The queries that actually surfaced this material, grouped by what they find. The vocabulary is the method: rollback material hides behind incident words, not feature words.

Incidents where the undo failed

  • deploy rollback postmortem "made it worse" OR "rollback failed"
  • "incident report" "roll back" "we decided" site:github.blog OR site:slack.engineering
  • outage "latent defect" "safe deployment" rings rollback
  • postmortem "we had never" rollback OR restore upgrade

Doctrine and design records

  • "ensuring rollback safety" OR "two-phase deployment" builders library
  • "bake time" OR "one-box" deployment pipeline rollback alarm
  • "push on green" rollout usenix login
  • "safe deployment" ring canary rollback paper NSDI OR OSDI OR ISSRE

The argument in the issue tracker

  • repo:kubernetes/kubernetes "automatic rollback" is:issue
  • repo:kubernetes/kubernetes rollback is:pr is:closed is:unmerged
  • "unable to automatically abort and rollback" argo OR flagger

The database fight

  • "down migrations" myth OR harmful OR "never run"
  • "expand contract" OR "expand/contract" schema rollback deploy
  • database rollback pitfalls "wasted time" OR "unknown state"
09

References

  1. U.S. SEC, In the Matter of Knight Capital Americas LLC, Release 34-70694 SEC, 2013-10-16. Checked 2026-10-05.
  2. CrowdStrike, Channel File 291 Incident Root Cause Analysis CrowdStrike, 2024-08-06. Checked 2026-10-05.
  3. John Graham-Cumming, Details of the Cloudflare outage on July 2, 2019 Cloudflare blog, 2019-07-12. Checked 2026-10-05.
  4. Reddit Infrastructure, You Broke Reddit: The Pi-Day Outage r/RedditEng, 2023-03-21. Checked 2026-10-05.
  5. GitHub, October 21 post-incident analysis GitHub blog, 2018-10-30. Checked 2026-10-05.
  6. BleepingComputer, Microsoft explains the cause of the recent Office 365 outage BleepingComputer, 2020-10. Checked 2026-10-05.
  7. The Register, Microsoft says a latent defect in its Safe Deployment Process downed Azure AD The Register, 2020-10-02. Checked 2026-10-05.
  8. Sandeep Pokkunuri, Ensuring rollback safety during deployments AWS Builders' Library, 2019. Checked 2026-10-05.
  9. Clare Liguori, Automating safe, hands-off deployments AWS Builders' Library (undated article; practice first presented at re:Invent 2017). Checked 2026-10-05.
  10. Beyer et al. (eds.), Site Reliability Engineering, chapter 1 Google / O'Reilly, 2016. Checked 2026-10-05.
  11. Klein, Betser, Monroe, Making "Push on Green" a Reality ;login: vol. 39 no. 5, USENIX, 2014-10. Checked 2026-10-05.
  12. Savor et al., Continuous Deployment at Facebook and OANDA ICSE 2016, ACM. Checked 2026-10-05.
  13. Li et al., Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment NSDI 2020, USENIX. Checked 2026-10-05.
  14. Meta, Making Deployments Safe at Meta: Health Checks for Continuous Change-Safety arXiv 2608.20513, accepted ISSRE 2026 industry track, 2026-08-20. Checked 2026-10-05.
  15. Kubernetes, issue #23211: Automatic rollback for failed deployments GitHub, opened 2016-03-18, closed unimplemented. Checked 2026-10-05.
  16. Kubernetes, website issue #55312: stale automatic-rollback promise in the Deployment docs GitHub, 2026-04-10, closed by PR #55402. Checked 2026-10-05.
  17. Kubernetes, PR #93807: Remove rollbackTo in deployment controller GitHub, opened 2020-08-08, approved, closed unmerged 2021-02-12. Checked 2026-10-05.
  18. Kubernetes, pkg/controller/deployment/rollback.go GitHub, master branch. Checked 2026-10-05.
  19. Argo project, Argo Rollouts GitHub README. Checked 2026-10-05.
  20. Flux CD, Flagger GitHub README. Checked 2026-10-05.
  21. Netflix, Automated Canary Analysis at Netflix with Kayenta Netflix Tech Blog, 2018-04. Checked 2026-10-05.
  22. Slack, Deploys at Slack Slack engineering blog, undated (circa 2020). Checked 2026-10-05.
  23. Sean McIlroy, The Scary Thing About Automating Deploys Slack engineering blog, 2024. Checked 2026-10-05.
  24. InfoQ, Slack Conquers Deployment Fears with Z-score Monitoring InfoQ, 2024-03. Checked 2026-10-05.
  25. getdx, How Slack fully automates deploys and anomaly detection with Z-scores getdx newsletter, 2024. Checked 2026-10-05.
  26. Ariga, The Myth of Down Migrations; Introducing Atlas Migrate Down atlasgo.io, 2024-04-01. Checked 2026-10-05.
  27. Octopus Deploy, Pitfalls with SQL rollbacks and automated database deployments Octopus blog, undated. Checked 2026-10-05.
  28. InfoQ podcast, Clare Liguori on Automating Safe and Hands-Off Deployments at AWS InfoQ, circa 2021. Checked 2026-10-05.
  29. Peter Ramensky, Amazon's approach to high-availability deployment (DOP404), slides AWS re:Invent 2019. Checked 2026-10-05.
  30. Cybersecurity Dive, CrowdStrike disruption direct losses to reach $5.4B for Fortune 500 Cybersecurity Dive, 2024-07. Checked 2026-10-05.
  31. Pragmatic Engineer, Interesting Learnings from Outages (#10) Pragmatic Engineer newsletter, 2023. Checked 2026-10-05.