In the Matter of Knight Capital Americas LLC (34-70694)
The regulator's reconstruction: code deployed to 7 of 8 servers with no second reviewer, a repurposed flag, and a rollback that activated dead code fleet-wide.
How production teams undo a bad change, reconstructed from six published incidents and the deployment doctrine of Amazon, Google, Meta, Netflix and Slack. What rollback actually is mechanically, the four ways it fails, and the decisions, made weeks before the incident, that keep the way back open.
The problem, who has solved it in production, and the one fact about rollback that the famous incidents keep re-teaching.
Strip the vocabulary away and the problem is this: a change you made is hurting production, and you must return the system to its last known-good behaviour faster than the damage compounds. The return trip is itself a change. It runs through the same machinery that just failed you, it is executed at the worst possible moment, under the least information and the most pressure, and it is almost always the least rehearsed operation the team owns.
The organisations that have written seriously about this problem solved it from both ends. Amazon's Builders' Library states the doctrine most plainly: every revision must be backwards compatible, meaning it "can be rolled back without errors or disruption", and that property is verified in the pipeline, not assumed, with an upgrade-then-rollback test run against part of the fleet before full deployment. Google's SRE book, Meta's health-check paper, Microsoft's Gandalf paper and Slack's ReleaseBot account all describe the same structure from different angles: expose the change progressively, watch an independent health signal, and make the return trip automatic because, in Google's words, humans in repetitive loops bring fatigue and inattention.
The fact the incidents keep re-teaching is the one this guide leads with. The most expensive deployment failure in the public record, Knight Capital, was not a failed rollback. The rollback executed cleanly. According to the SEC's order, engineers uninstalled the new code from the seven servers where it had been deployed correctly, and because the new code had repurposed an old feature flag, the rollback re-activated eight-year-old dead code on all eight servers instead of one. Rollback restores your code. It does not restore the context your code runs in: the flags, the configuration, the data, the traffic. Every failure class in section 4 is a variation on that asymmetry.
Scope, stated up front: this guide covers getting a bad application deploy, config push or content update out of production, and the schema and state contracts that make that possible. It deliberately does not cover restoring data from backups (see the sibling guide on proving you can restore), the mechanics of live schema migration (covered in changing the schema under load), or recalling software from devices you do not control (updating software you cannot recall). Incident process in general is in learning from production failure.
The shape that recurs across Amazon, Meta, Azure, Netflix and Slack, and the mechanical fact about rollback that explains most of the failures.
Start with the mechanism, because it is widely misunderstood. There is no reverse gear.
In every system with a published implementation, rolling back means deploying the previous
version forward. The Kubernetes deployment controller's rollback.go literally
"compares the templates of the provided deployment and replica set and updates the
deployment with the replica set template in case they are different": the old pod template
is copied onto the Deployment and an ordinary rollout proceeds. Slack's account of ReleaseBot
says the same thing from the operations side; per the getdx interview with Sean McIlroy,
rollback is Slack's fastest deploy, taking a minute or two, precisely because it is a deploy
with the build and test steps already paid for.
That one fact carries three consequences. First, everything that makes your deploys slow makes your rollbacks slow; a 40-minute pipeline is a 40-minute undo. Second, rollback inherits every dependency of the forward path, so when the forward machinery is what broke, the way back is blocked too: CrowdStrike's revert shipped 78 minutes after the bad content, but hosts that had already loaded it were boot-looping and could not fetch anything, and Azure AD's automated rollback in September 2020 failed because the same defect that mis-targeted the deployment had corrupted the metadata the rollback needed. Third, whether the old version still works is not a property of the rollback tool at all. It is a property of the compatibility contract between versions, which was fixed weeks earlier, when the change was designed. Amazon's doctrine makes this explicit: rollback safety is decided at design time and verified in pre-production, because by the time you need it, it is too late to add.
Amazon's one-box stage serves at most ten percent of a zone's requests before wave deployment begins, per Clare Liguori's account. Slack sends builds to a dogfood tier, then a canary taking about 2% of production traffic. Azure deploys in rings with bake time between them. The point is not caution for its own sake; it is that the undo cost in figure 1 is proportional to exposure at detection time.
Every implementation separates the thing that deploys from the thing that judges. Slack's ReleaseBot watches z-scores on error and latency series. Kayenta compares canary and baseline statistically and fails the canary on significant degradation. Meta's Service Health Checker runs templated metric checks wired into every phased rollout. Gandalf goes furthest: an ensemble that attributes fault signals to one of many concurrent rollouts before deciding no-go.
Sources: InfoQ on Slack, Netflix, Meta, Microsoft
Azure's deployment engine subscribes to Gandalf's go/no-go decisions and stops a rollout on no-go. AWS pipelines roll back the one-box automatically without promoting further. The executor itself is boring by design: deploy the previous artifact, shift traffic back, or both. Boring is the feature. The executor runs on the worst day, so every clever branch in it is a liability.
Sources: Gandalf, NSDI 2020, AWS
The load-bearing component has no runtime presence at all. Amazon requires serialized data, APIs and stored state to be readable by version N when written by version N+1, enforced through two-phase deployments (figure 3) and verified by deploying to half a fleet, rolling back, and checking both directions. Without this contract the executor is a loaded gun: the artifact reverts, the data does not.
Source: AWS Builders' Library, 2019
Where implementations diverge is instructive. Who decides varies from fully automatic (AWS pipelines, Gandalf, ReleaseBot after 2024) to advisory (Kayenta can route to a human approval path) to fully manual (Reddit's operators, on the day). What reverts first varies: Facebook and OANDA, per Savor et al., reach for config kill switches before artifact rollbacks, while Cloudflare's 2019 recovery was a global kill switch on the WAF ruleset, not a redeploy. And where the capability lives varies most of all, which is section 3's second decision.
Four forks, each with the condition that flips the answer. The pattern underneath all four: rollback is a capability you buy at design time, not a button you press at incident time.
rollbackTo) was
deprecated in favour of a client-side kubectl rollout undo.| Decision | Chosen | Rejected | Because | Evidence |
|---|---|---|---|---|
| First move | Roll back, reflexively | Fix forward under pressure | The incident is the worst time to diagnose | SRE book, Slack |
| Trigger | Automated, metric-gated | Human per-deploy vigilance | Fatigue; detection precision now supports it | Gandalf, Meta SHC |
| Where it lives | Layer above the orchestrator | Auto-rollback in Kubernetes core | Core keeps forward-only primitives; policy varies too much | #23211, Argo |
| Schema | Forward-only + expand/contract | Down migrations in production | Asymmetry: reversing an additive change destroys data | Atlas, AWS |
| Exposure unit | One-box / canary first | Fleet-wide push | Undo cost is proportional to exposure at detection | Liguori, Slack |
Every decision above reduces to one question: how old is the newest thing the previous version cannot understand? If the answer is "nothing, by contract" you may roll back blind, automatically, in minutes. The moment flags, config, schema or serialized data escape that contract, the undo stops being an operation and becomes a diagnosis.
Six published incidents, four failure classes. None of them is a rollback tool malfunctioning; every one is an assumption about the way back that nobody had tested.
Grouping the public record, four classes cover it. Class 1: rollback restores code, not context, the Knight asymmetry. Class 2: the undo nobody rehearsed, where the rollback was possible but the team could not price its risk, so the decision itself became the outage. Class 3: the undo needs the system to be alive, where the rollback depends on exactly the machinery the failure took out. Class 4: state closes the window, where every minute of forward operation makes the return trip more expensive, until it is no longer a rollback but a reconciliation project.
What undo speed, exposure and failure actually measure out to in the published record.
| Metric | Value | At | Context | As of | Source |
|---|---|---|---|---|---|
| Outages caused by live-system changes | ~70% | Stated in the SRE book's change-management discussion; internal, undated measurement | 2016 | SRE book | |
| Rollback execution time | 1-2 min | Slack | ReleaseBot's rollback, described as the fastest deploy they run | 2024 | getdx interview |
| Canary exposure before promotion | ~2% | Slack | Share of production traffic on the canary tier | 2020 | Deploys at Slack |
| One-box exposure cap | 10% | Amazon | Maximum share of requests served by the one-box stage | 2021 | Liguori |
| No-go detection precision | 92.4% / 100% recall | Azure | Gandalf, data-plane rollouts over 18 months in production | 2020 | NSDI 2020 |
| Bad-deploy loss rate | $460M / 45 min | Knight | Pre-tax loss between flag-on and full stop, rollback included | 2012 | SEC order |
| Push-to-revert latency | 78 min | CrowdStrike | 04:09 UTC push to 05:27 UTC revert; derived from the RCA timeline | 2024 | RCA |
| Hosts crashed before revert landed | ~8.5M | Windows fleet | Microsoft's estimate, under 1% of Windows machines | 2024 | Cybersecurity Dive |
| Modelled Fortune 500 direct losses | $5.4B | Parametrix | Insurer's model, average $44M per affected company; claimed, not audited | 2024 | Cybersecurity Dive |
| Decision latency dominating an outage | 314 min total | Hours of it spent deciding whether the untested rollback was safe | 2023 | postmortem | |
| Divergence that closed the failback window | ~40 min of writes | GitHub | Unreplicated writes on both sides after a 43-second partition; 24h11m to full recovery | 2018 | GitHub |
| Deploy frequency under this discipline | ~12/day | Slack | Scheduled webapp deploys, each a potential rollback rehearsal | 2020 | Deploys at Slack |
Measured: the SEC, CrowdStrike, GitHub, Reddit and Cloudflare figures come from primary incident documents. Self-reported: Gandalf's precision and recall, Slack's rollback time and Amazon's exposure caps are the operators describing their own systems, with no independent measurement published. Modelled: Parametrix's $5.4B is an insurer's estimate and should be quoted as such. Undated: Google's 70% has no published methodology or date; treat it as directional, not as a planning constant.
Every source behind this page, graded. Filter by kind. Links were retrieved 2026-10-05; this build's container could re-verify GitHub-hosted links directly, others via server-side retrieval only.
The regulator's reconstruction: code deployed to 7 of 8 servers with no second reviewer, a repurposed flag, and a rollback that activated dead code fleet-wide.
21 fields defined, 20 supplied; the revert shipped in 78 minutes and could not reach hosts that no longer booted. Remediations are all rollout-shaped: rings, validation, customer control.
A WAF rule skipped staged rollout by design and saturated every CPU on the edge. Recovery came from a global kill switch, not a redeploy.
A Kubernetes upgrade deleted the node label years-old unversioned config depended on. The restore worked; pricing an untested restore mid-incident cost hours.
43 seconds of partition, 40 minutes of diverged writes, and the failback was gone. GitHub chose integrity over recovery time and paid 24 hours for it, explicitly.
One latent defect mis-targeted all rings at once and corrupted the metadata the automated rollback needed. Cited through BleepingComputer and The Register, which quote the RCA; Microsoft's own status-history page is not stably linkable.
Amazon's doctrine: backwards compatibility at every revision, the two-phase deployment for serialized data, and the upgrade-rollback test as a pipeline gate.
The pipeline anatomy: one-box capped at 10% of requests, wave deployment, bake time, alarm-gated promotion, automatic rollback with no human in the loop.
The feature request, probe-driven auto-undo with a failure budget, parked in the SIG Apps backlog and closed without implementation.
The docs said the controller would roll back automatically "in the future, once automatic rollback will be implemented". A decade later the sentence was removed rather than the feature added (PR #55402).
The cleanup removing dead server-side rollback logic was approved in two days, went stale waiting on a rebase, and the bot closed it. The deprecated code path outlived the attempt to delete it.
The implementation: rollback copies the old ReplicaSet's template onto the Deployment and rolls forward; the old API survives only as a deprecated annotation that is cleared after use.
States the gap it fills in one line: a Deployment "can halt the progression, but unable to automatically abort and rollback the update". Features: automated rollbacks and promotions, metric analysis.
Progressive delivery operator: shifts traffic gradually while measuring metrics and running conformance tests, reverting on threshold breach.
The 70% figure and the prescription: progressive rollouts, fast accurate detection, safe rollback, humans out of the repetitive loop.
The insurer's model: $5.4B direct losses, $44M average per affected company, 10-20% insured. A modelled estimate, labelled as such here.
Twelve scheduled deploys a day, a rotating deploy commander, dogfood then a 2% canary, rollback and hotfix as the two standing remedies.
ReleaseBot replaced the human deploy commander: z-score anomaly detection, automatic pause and rollback. The getdx interview adds the number: rollback is their fastest deploy, one to two minutes.
Statistical comparison of canary and baseline metrics; on significant degradation the canary is aborted and traffic routes back to stable. Built with Google, open-sourced into Spinnaker.
The argument that symmetric down files cannot deliver what they promise: failed migrations leave unknown states, and reversing additive changes destroys data.
Independent close reading of the Pi-Day postmortem, preserving the line this guide leans on: hours to decide that a rollback, a high-risk action on its own, was the best course.
The decay argument from a deployment vendor: rollback scripts that are never run are wasted work, and the motivation to write them falls with every successful deploy.
The field study: remedies for a bad deploy in practice are rollback, rapid hotfix, or a config change disabling the feature, in that menu, across two very different firms.
Azure's go/no-go analytics: ensemble fault attribution across concurrent rollouts, 92.4% precision and 100% recall on the data plane over 18 months, wired directly into the deployment engine.
The Service Health Checker: templated metric checks composed by service owners, integrated with tiered and phased rollouts to trigger automatic rollback on regressions, and a frank account of noise, drift and alert fatigue at fleet scale.
Google's early public account of mechanised rollouts: if the tests are good and the build is good, push, with rollback as a first-class, rehearsed step rather than an escape hatch.
The spoken companion to the Builders' Library article: one-box stages, alarm-gated waves, automatic rollback as the default behaviour of every pipeline.
The conference treatment of fractional deployments, traffic shifting, bake time and rollback. Cited from the published deck; this build's environment could not retrieve the video for timestamped claims.
Seven rungs from a toy blue-green flip to breaking your own rollback machinery on purpose. The crossing from toy to real is rung four.
Run v1 and v2 of a trivial service behind nginx or Caddy; flip traffic between them with a config reload. Script the flip both ways.
Done when: either direction completes in under 5 seconds with zero dropped requests under light load. Teaches: traffic shift as the cheapest undo.
Ship a deliberately broken v2 (10% error rate). A watchdog polls the error ratio and flips traffic back without you.
Done when: injected failure to full recovery is under 60 seconds, unattended, ten runs in a row. Teaches: the gate-and-executor loop, and your first false-positive tuning.
Have v2 write a new field into a queue message or cache entry that v1 cannot parse. Roll back and watch v1 crash on v2's leftovers.
Done when: you can reproduce the post-rollback failure on demand and explain exactly which artifact carried the poison. Teaches: the Knight asymmetry at toy scale.
Re-ship the same change as two releases: first teach the reader to accept both formats, then start writing the new one. Roll back from each phase.
Done when: rollback from every intermediate state leaves the system healthy, demonstrated by test. Teaches: the compatibility contract that underwrites every real rollback.
Add a column the new code uses while the old code still runs, then contract only after the rollback window closes. No down migration anywhere.
Done when: v1 runs correctly against the fully migrated schema, proven in CI. Teaches: why forward-only migration is a rollback feature, not a limitation.
Amazon-style gate: CI deploys the new build to half your toy fleet, rolls it back, and asserts health in both directions before anything promotes.
Done when: a commit that breaks backward compatibility of stored data is rejected by the pipeline, not by a human. Teaches: rollback safety as a testable property.
Game-day your own machinery: corrupt the artifact store, kill the config service mid-rollback, make the previous image unpullable. Measure time to last known good.
Done when: you have a measured number for "time to previous version with the happy path blocked" and a written fallback. Teaches: the Azure AD lesson; the guardrail is also software.
The queries that actually surfaced this material, grouped by what they find. The vocabulary is the method: rollback material hides behind incident words, not feature words.
deploy rollback postmortem "made it worse" OR "rollback failed""incident report" "roll back" "we decided" site:github.blog OR site:slack.engineeringoutage "latent defect" "safe deployment" rings rollbackpostmortem "we had never" rollback OR restore upgrade"ensuring rollback safety" OR "two-phase deployment" builders library"bake time" OR "one-box" deployment pipeline rollback alarm"push on green" rollout usenix login"safe deployment" ring canary rollback paper NSDI OR OSDI OR ISSRErepo:kubernetes/kubernetes "automatic rollback" is:issuerepo:kubernetes/kubernetes rollback is:pr is:closed is:unmerged"unable to automatically abort and rollback" argo OR flagger"down migrations" myth OR harmful OR "never run""expand contract" OR "expand/contract" schema rollback deploydatabase rollback pitfalls "wasted time" OR "unknown state"