Evidence ledger 28 sources Checked 05 Oct 2026

Evidence ledger

One row per claim in The way back is also a deploy: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

Checked 2026-10-05. Research constraint for this run: the build container's network policy allowed direct fetches only to github.com; all other sources were retrieved through the server-side web search tool, which returns live page content but not the full page. Quotes below were copied from retrieved content in this session. Link liveness could not be re-verified from inside the container (verifier reports them as unreachable warnings, not dead links).

# Claim Org / author Tier Published Checked URL Supporting quote / figure
1 Knight's attempted rollback amplified the failure: uninstalling the new RLP code from the seven correct servers re-activated dormant Power Peg code on all eight, because the new code had repurposed Power Peg's activation flag U.S. SEC (In the Matter of Knight Capital Americas LLC, Release 34-70694) postmortem 2013-10-16 2026-10-05 https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf "Knight uninstalled the new RLP code from the seven servers where it had been deployed correctly"; the repurposed flag then triggered Power Peg on all eight servers
2 The triggering defect was a deployment gap: a technician did not copy the new code to one of the eight SMARS servers and no second technician reviewed the deployment U.S. SEC postmortem 2013-10-16 2026-10-05 https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf "did not copy the new code to one of the eight SMARS computer servers… no second technician review"
3 Loss figure: ~$460M pre-tax loss over ~45 minutes of trading on 2012-08-01 U.S. SEC / widely corroborated postmortem 2013-10-16 2026-10-05 https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf $460 million loss, 45 minutes
4 CrowdStrike Channel File 291: faulty content pushed 04:09 UTC 2024-07-19, reverted 05:27 UTC; machines that had loaded it kept crashing because the revert could not reach hosts that could no longer boot CrowdStrike external RCA postmortem 2024-08-06 2026-10-05 https://www.crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf pushed 04:09 UTC, reverted 05:27 UTC; out-of-bounds read → boot loop
5 CrowdStrike remediation adopted staged rollout of content updates (canary then rings), i.e. the rollback lesson became a rollout lesson CrowdStrike external RCA postmortem 2024-08-06 2026-10-05 https://www.crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf RCA mitigations include staged deployment of content, customer control over rollout
6 Cloudflare 2019-07-02: one WAF rule deployed globally with no gradual rollout dropped ~82% of traffic for 27 minutes (13:42–14:09 UTC); recovery was a 'global kill' switch on the WAF managed rulesets, not a redeploy Cloudflare (via contemporaneous accounts of the official postmortem) postmortem 2019-07-12 2026-10-05 https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019/ 13:42–14:09 UTC; "global kill" at 14:02 decision, traffic restored 14:09; CPU protection had been removed in an earlier refactor
7 Reddit Pi-Day 2023: 314-minute outage after a Kubernetes 1.23→1.24 upgrade; the team took hours to decide on restore because rollback of a cluster upgrade was itself "a high-risk action" they had not practised Reddit infrastructure team (r/RedditEng postmortem) postmortem 2023-03-21 2026-10-05 https://www.reddit.com/r/RedditEng/comments/11xx5o0/you_broke_reddit_the_piday_outage/ 314 minutes down; root cause: 1.24 removed the master node label, Calico route reflectors selected on it; "It took hours to decide that a rollback, a high-risk action on its own, was the best course of action" (as reported by Pragmatic Engineer's account of the postmortem)
8 Amazon's deployment doctrine: every revision must be backwards compatible so it can be rolled back "without errors or disruption"; rollback safety is tested by an upgrade-rollback test that deploys to half the fleet and rolls back AWS Builders' Library — Sandeep Pokkunuri, "Ensuring rollback safety during deployments" adr 2019-12 2026-10-05 https://aws.amazon.com/builders-library/ensuring-rollback-safety-during-deployments/ "A version of software that can be rolled back without errors or disruption… is called backwards compatible"; two-phase deployment technique for serialized data
9 Kubernetes has repeatedly declined to build automatic rollback into Deployments; the feature request sat in SIG Apps backlog and was closed Kubernetes (issue #23211, "Automatic rollback for failed deployments") source 2016-03 (opened) 2026-10-05 https://github.com/kubernetes/kubernetes/issues/23211 Closed; proposal for probe-driven auto-undo with maxFailures; assigned SIG Apps backlog priority
10 A Kubernetes rollback is implemented as a forward deploy: rollbackToTemplate() copies the old ReplicaSet's pod template onto the Deployment and the rollout proceeds normally; the deprecated rollbackTo server-side API survives only as an annotation (getRollbackTo/setRollbackTo) Kubernetes source, pkg/controller/deployment/rollback.go source current master 2026-10-05 https://github.com/kubernetes/kubernetes/blob/master/pkg/controller/deployment/rollback.go "compares the templates of the provided deployment and replica set and updates the deployment with the replica set template in case they are different"
11 The Deployment docs promised automatic rollback "in the future" for years; issue filed 2026-04-10 notes it was never implemented and has no plans, closed by docs PR #55402 removing the sentence Kubernetes (website issue #55312) source 2026-04-10 2026-10-05 https://github.com/kubernetes/website/issues/55312 Doc sentence: "In the future, once automatic rollback will be implemented, the Deployment controller will roll back a Deployment as soon as it observes such a condition."
12 PR removing the dead rollbackTo logic was approved (janetkuo, 2020-08-10) yet closed unmerged 2021-02-12 by the staleness bot after a rebase was never done Kubernetes (PR #93807) source opened 2020-08-08 2026-10-05 https://github.com/kubernetes/kubernetes/pull/93807 "Rotten issues close after 30d of inactivity"; approved but never merged
13 Argo Rollouts exists because "Can halt the progression, but unable to automatically abort and rollback the update" is a stated Kubernetes Deployment limitation; features include "Automated rollbacks and promotions" Argo project README source current 2026-10-05 https://github.com/argoproj/argo-rollouts quotes as shown
14 Flagger frames progressive delivery as risk reduction: "It reduces the risk of introducing a new software version in production by gradually shifting traffic to the new version while measuring metrics and running conformance tests" Flux CD / Flagger README source current 2026-10-05 https://github.com/fluxcd/flagger quote as shown
15 Google SRE: "roughly 70% of outages are due to changes in a live system"; recommended mitigations are progressive rollouts, fast accurate detection, and safe rollback, with humans removed from the loop Google SRE book, ch. 1 (Change management) casestudy 2016 2026-10-05 https://sre.google/sre-book/introduction/ quote as shown
16 GitHub 2018-10-21: after a 43-second network partition, the West Coast cluster had ingested ~40 minutes of writes; with unreplicated writes on both sides GitHub could not fail back safely and chose data integrity over time-to-recovery; degraded for 24h11m GitHub post-incident analysis postmortem 2018-10-30 2026-10-05 https://github.blog/news-insights/company-news/oct21-post-incident-analysis/ "database clusters in both data centers now contained writes that were not present in the other"; 24 hours 11 minutes
17 Azure AD 2020-09-28: an update meant for a validation test ring deployed to all rings at once because of a latent SDP defect; the automated rollback then failed because the same defect corrupted deployment metadata, forcing manual rollback; ~5h auth outage Microsoft RCA as quoted by press (BleepingComputer; The Register) postmortem 2020-10-01 2026-10-05 https://www.bleepingcomputer.com/news/microsoft/microsoft-explains-the-cause-of-the-recent-office-365-outage/ "the SDP system failed to correctly target the validation test ring... all rings were targeted concurrently"; rollback "required a much longer manual rollback" (also https://www.theregister.com/2020/10/02/microsoft_azure_bug/)
18 Slack deploys ~12 times/day behind a deploy commander; rollout goes dogfood then canary (~2% of traffic) Slack engineering, "Deploys at Slack" blog undated (checked) 2026-10-05 https://slack.engineering/deploys-at-slack/ 12 scheduled deploys/day; canary ~2% of production traffic
19 Slack's ReleaseBot automated the deploy commander role using z-score anomaly detection and can pause or roll back on its own; rollback is Slack's fastest deploy, one-two minutes Slack engineering, "The Scary Thing About Automating Deploys" (+ InfoQ summary, getdx interview) blog 2024 (InfoQ 2024-03) 2026-10-05 https://slack.engineering/the-scary-thing-about-automating-deploys/ "ReleaseBot's rollback is instantaneous and is their fastest deploy, taking just a minute or two" (getdx account); https://www.infoq.com/news/2024/03/slack-z-score-monitoring/
20 Kayenta compares canary vs baseline metrics statistically and can automatically fail the canary, routing all traffic back to stable Netflix Tech Blog (Kayenta, with Google) blog 2018-04 2026-10-05 https://netflixtechblog.com/automated-canary-analysis-at-netflix-with-kayenta-3260bc7acc69 "if significant degradation is detected, the canary is aborted and all traffic is routed to the stable version"
21 Down-migration files are argued to be a liability: a failed upgrade leaves the DB in a state the down file does not expect, and reversing an additive change destroys data (DROP COLUMN deletes the column's data) Ariga (Atlas), "The Myth of Down Migrations" blog 2024-04-01 2026-10-05 https://atlasgo.io/blog/2024/04/01/migrate-down "If a migration fails, the database might be in an unknown state"; inverse of ADD COLUMN "deletes all the data in that column"
22 Rollback scripts that are never exercised decay: "If the script is never run, then it turns out to be wasted time"; motivation to write them decreases with each successful deploy Octopus Deploy vendor undated (checked) 2026-10-05 https://octopus.com/blog/database-rollbacks-pitfalls quote as shown
23 Facebook/OANDA remedies for a bad deploy, in practice: revert via deployment rollback, rapid hotfix, or config change disabling the feature; scale evidence: 20x engineers, 50x code, flat productivity Savor et al., ICSE 2016 paper 2016-05 2026-10-05 https://dl.acm.org/doi/10.1145/2889160.2889223 "remedies include reverting the deployed update through a deployment rollback, rapid deployment of a hotfix, or configuration changes to disable the problematic feature"
24 Azure's deployment engine subscribes to Gandalf go/no-go decisions; 92.4% precision / 100% recall on data-plane rollouts (no high-impact Azure Compute outages caused by bad rollouts in the measured period); 94.9%/99.8% control plane; >20TB/day analyzed Li et al., NSDI 2020 (Gandalf) paper 2020-02 2026-10-05 https://www.usenix.org/system/files/nsdi20spring_li_prepub.pdf figures as shown
25 Meta's Service Health Checker integrates with tiered and phased rollouts "to trigger automatic rollback on regressions" across thousands of services; operational problems at scale: noise, alert fatigue, drift Meta (ISSRE 2026 industry track, arXiv:2608.20513) paper 2026-08-20 2026-10-05 https://arxiv.org/abs/2608.20513 abstract wording as shown
26 Google "Push on Green": mechanized rollout best practice defined as "if the tests are good, the build is good, go push it", with rollback as a first-class step Klein, Betser, Monroe, ;login: 39(5) paper 2014-10 2026-10-05 https://www.usenix.org/system/files/login/articles/login_1410_05_klein.pdf quote as shown
27 Amazon one-box stage serves at most ten percent of requests; on negative impact the pipeline rolls back automatically without promoting further Clare Liguori (AWS), InfoQ podcast + Builders' Library article talk article undated; podcast 2021 era 2026-10-05 https://www.infoq.com/podcasts/clare-liguori-aws-deployments/ "a single box serves at most ten percent of overall requests"; auto-rollback without human intervention (also https://builder.aws.com/content/3ErTKQOTKc5NIw031UePBPxTQ6I/automating-safe-hands-off-deployments)
28 Amazon re:Invent 2019 DOP404 covers fractional deployments, rollback, traffic shifting and preproduction strategies for high-availability deployment Peter Ramensky (AWS), re:Invent 2019 deck talk 2019-12 2026-10-05 https://d1.awsstatic.com/events/reinvent/2019/REPEAT_1_Amazon%27s_approach_to_high-availability_deployment_DOP404-R1.pdf.pdf deck exists; topics per listing. Timestamped video claims unavailable from this container (network policy); accepted limitation
29 CrowdStrike impact scale: ~8.5M Windows devices (Microsoft estimate); Parametrix puts Fortune 500 direct losses at $5.4B, average $44M per affected company Microsoft / Parametrix via Cybersecurity Dive casestudy 2024-07 2026-10-05 https://www.cybersecuritydive.com/news/crowdstrike-cost-fortune-500-losses-cyber-insurance/722396/ figures as shown
30 Reddit's own wording on the restore decision and the untested path; corroborating secondary account Pragmatic Engineer newsletter analysis of the Reddit postmortem blog 2023-04 2026-10-05 https://newsletter.pragmaticengineer.com/p/real-world-engineering-10 "It took hours to decide that a rollback, a high-risk action on its own, was the best course of action"