Evidence ledger
One row per claim in Letting humans into production: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Guide: Letting humans into production: zero standing privilege, just-in-time access, and the
break-glass path, reconstructed from six breach reports and five builder accounts
Category: security-and-identity · Research date: 2026-09-28 · All sources located and checked 2026-09-28.
How each source was read, stated up front. This session's network egress policy resolved
github.com, raw.githubusercontent.com, gitlab.com and cloud.google.com directly and
refused every other host at the proxy. Sources marked [direct] below were fetched in full
in this session (the Building Secure and Reliable Systems chapter from its source repository,
the GitLab runbooks, handbook and infrastructure tracker from gitlab.com, the ConsoleMe README
and login.gov handbook from raw.githubusercontent.com). Sources marked [search] could not be
fetched whole; their existence, URL, date and the quoted material were confirmed through the
session's server-side web search, which returns live excerpts from the page, and the quotes
recorded here are strings that search surfaced verbatim. Where only a paraphrase could be
confirmed, the right-hand column says "reported as" instead of quoting. No URL below was cited
from memory; every one was returned by a live search or fetch in this session. The page's link
verifier will report the blocked hosts as unreachable from this container; that is the proxy,
not the links.
One row per claim. The right-hand column is copied text, not paraphrase, except where marked.
| # | Org | Title | Tier | Published | Checked | URL | Claim taken from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | Building Secure and Reliable Systems, Ch. 5 "Design for Least Privilege" (Barrett, Joyner, Ward) [direct] | casestudy | 2020 | 2026-09-28 | https://google.github.io/building-secure-and-reliable-systems/raw/ch05.html | Zero Touch interfaces exist to remove direct human access to production, for reliability as much as security. | "make Google safer and reduce outages by removing direct human access to production roles. Instead, humans have indirect access to production through tooling and automation that make predictable and controlled changes to production infrastructure." | |
| 2 | Same chapter, breakglass section [direct] | casestudy | 2020 | 2026-09-28 | https://google.github.io/building-secure-and-reliable-systems/raw/ch05.html | Breakglass is defined as a total bypass of authorization, not a faster approval. | "a breakglass mechanism provides access to your system in an emergency situation and bypasses your authorization system completely." | |
| 3 | Same chapter, auditing culture [direct] | casestudy | 2020 | 2026-09-28 | https://google.github.io/building-secure-and-reliable-systems/raw/ch05.html | The failure mode of breakglass is normalisation, and the control against it is cultural, not technical. | "Without cultural reinforcement, audits can become rubber stamps, and breakglass use can become an everyday occurrence, losing its sense of importance or urgency." | |
| 4 | Same chapter, three-factor authorization [direct] | casestudy | 2020 | 2026-09-28 | https://google.github.io/building-secure-and-reliable-systems/raw/ch05.html | Multi-party approval fails as a control when every approver works from the same managed fleet an attacker has compromised. | "In a large organization, MPA often has one key weakness that can be exploited by a determined and persistent attacker: all of the 'multiple parties' use the same centrally managed workstations." | |
| 5 | Google Cloud | "How Google protects its production services" [search] | vendor | current, checked 2026 | 2026-09-28 | https://docs.cloud.google.com/docs/security/production-services-protection | Google attributes a measured share of its outages to changes ZTP would have blocked, and forbids unilateral access to foundational services even in emergencies. | Reported as: Google estimates ~13% of all Google-evaluated outages could have been prevented or mitigated with Zero Touch Prod; "Google doesn't allow unilateral access to foundational services — even emergency access requires approval from other Google personnel." |
| 6 | Google / USENIX | "Zero Touch Prod: Towards Safer and More Secure Production Environments", SREcon19 EMEA (Czapiński, Wolafka) [search] | talk | 2019-10 | 2026-09-28 | https://www.usenix.org/conference/srecon19emea/presentation/czapinski | The canonical ZTP rule: three legal paths into production, and only three. | Abstract: "Zero Touch Prod requires every change in production to be made by automation (instead of humans), prevalidated by software, or triggered through an audited breakglass mechanism." |
| 7 | NYDFS (regulator) | Twitter Investigation Report, October 2020 [search] | postmortem | 2020-10-14 | 2026-09-28 | https://www.dfs.ny.gov/reports_and_publications/press_releases/twitter_report (PDF: https://www.dfs.ny.gov/system/files/documents/2026/07/Twitter-Investigation-Report.pdf) | Standing access at population scale: over 1,000 people held the internal account tools; entry was a phone call impersonating the IT help desk about VPN problems; the haul was at least $118,000 in bitcoin; Twitter had no CISO since December 2019. | Reported as: "over 1,000 Twitter employees still had access" to internal tools; hackers "claimed to be calling from the Help Desk in Twitter's IT department" about VPN issues; scam "netted the teen hackers at least $118,000"; Twitter "had not had a chief information security officer (CISO) since December 2019." |
| 8 | CISA / CSRB (regulator) | Review of the Summer 2023 Microsoft Exchange Online Intrusion [search] | postmortem | 2024-03-20 | 2026-09-28 | https://www.cisa.gov/sites/default/files/2025-03/CSRBReviewOfTheSummer2023MEOIntrusion508.pdf | A 2016 signing key that was never rotated let Storm-0558 read the mail of 22 organisations and 500+ individuals; Microsoft could not establish how the key left its custody. | "cascade of Microsoft's avoidable errors"; "Microsoft's security culture was inadequate and requires an overhaul"; compromised "Microsoft Exchange Online mailboxes of 22 organizations and over 500 individuals"; the consumer signing key "had not been rotated since 2016"; "Microsoft has no evidence or logs showing the stolen key's presence in or exfiltration from a crash dump." |
| 9 | Cloudflare | "Thanksgiving 2023 security incident" [search] | postmortem | 2024-02-01 | 2026-09-28 | https://blog.cloudflare.com/thanksgiving-2023-security-incident | Four service credentials missed in a rotation of thousands, kept because they were "believed unused", admitted a nation-state actor to the internal Atlassian estate. | "we failed to rotate one service token and three service accounts (out of thousands) of credentials"; the four: a Moveworks service token, a Smartsheet service account with administrative access to Jira, a Bitbucket service account, an AWS account; believed "they were unused." |
| 10 | Cloudflare | Same post, remediation figures [search] | postmortem | 2024-02-01 | 2026-09-28 | https://blog.cloudflare.com/thanksgiving-2023-security-incident | The cost of losing credential inventory certainty is a company-wide reset. | Rotated ~5,000 production credentials; forensic triage of 4,893 systems; reimaged and rebooted every machine in the global network. |
| 11 | CircleCI | "CircleCI incident report for January 4, 2023 security incident" [search] | postmortem | 2023-01-13 | 2026-09-28 | https://circleci.com/blog/jan-4-2023-incident-report/ | Malware on one engineer's laptop stole a live, MFA-backed SSO session; the engineer was targeted because their role could mint production access tokens; dwell time from infection (Dec 16) to understanding (Jan 4) was 19 days, and the alert came from a customer. | Reported as: malware undetected by antivirus was "able to execute session cookie theft", letting attackers steal "a valid, 2FA-backed SSO session" and "escalate access to a subset of the company's production systems" via "the elevated permissions granted to the targeted employee"; reconnaissance Dec 19, exfiltration Dec 22, customer report of suspicious GitHub OAuth activity Dec 29. |
| 12 | Okta | "Unauthorized Access to Okta's Support Case Management System: Root Cause and Remediation" (D. Bradbury) [search] | postmortem | 2023-11-03 | 2026-09-28 | https://sec.okta.com/articles/2023/11/unauthorized-access-oktas-support-case-management-system-root-cause/ | The credential that opened the support system was a service account whose password had been saved into an employee's personal Google profile on a managed laptop; 134 customers' files were exposed over a 20-day window. | Reported as: access "leveraged a service account stored in the system itself" with permission to view and update support cases; the employee "had signed in to their personal Google profile on the Chrome browser of their Okta-managed laptop", where the service account's username and password had been saved; Sept 28 to Oct 17, 2023; 134 customers (under 1%). |
| 13 | BeyondTrust | "BeyondTrust Discovers Breach of Okta Support Unit" [search] | postmortem | 2023-10-20 | 2026-09-28 | https://www.beyondtrust.com/blog/entry/okta-support-unit-breach | The customer that watched its own identity plane caught the vendor-side compromise in half an hour; the vendor took ~two weeks to confirm. | Reported as: BeyondTrust's Okta administrator uploaded a HAR file on Oct 2 and the team "detected suspicious activity involving the session cookie within 30 minutes of sharing the file", flagging "an identity-centric attack on an in-house Okta administrator account", with no impact to its own infrastructure. |
| 14 | Uber | "Security update", Uber Newsroom, September 2022 | vendor | 2022-09-16 | 2026-09-28 | (URL not cited) | Uber's sanctioned account stated no evidence of access to sensitive user data. This session's web search could not confirm the exact newsroom URL, so it is NOT cited in the page; the statement is carried by the GitGuardian source (row 15), which reports it. | "We have no evidence that the incident involved access to sensitive user data (like trip history)." (as reported by row 15) |
| 15 | GitGuardian | "Uber Breach 2022 – Everything You Need to Know" [search] | blog | 2022-09 (updated) | 2026-09-28 | https://blog.gitguardian.com/uber-breach-2022/ | The privilege-management system itself was the pivot: a network share held a PowerShell script with hardcoded admin credentials for the PAM tool, whose vault opened every downstream system. | Reported as: after MFA-fatigue plus a WhatsApp call posing as Uber IT won VPN access, the attacker found PowerShell scripts on a network share, "one of which contained hardcoded credentials for a domain admin account for Thycotic, Uber's Privileged Access Management (PAM) solution", yielding admin over AWS, GCP, Google Drive, Slack, SentinelOne, HackerOne and internal dashboards. |
| 16 | Mercari | "Shifting to Zero Touch Production" [search] | blog | 2022-01-26 | 2026-09-28 | https://engineering.mercari.com/en/blog/entry/20220126-shifting-to-zero-touch-production/ | A second organisation independently adopted the ZTP triad and stated the rule for manual changes. | "an approval or audited break glass system should be used for manual changes." |
| 17 | Mercari | "Promote Zero Touch Production – further features of Carrier" [search] | blog | 2022-02-01 | 2026-09-28 | https://engineering.mercari.com/en/blog/entry/20220201-promote-zero-touch-production-further-features-of-carrier/ | Mercari's in-house gateway makes breakglass approval-free but loud: every use alarms a Slack channel and is audited. | "BreakGlass is a way for developers to get permissions without reviewer approval in emergency situations"; "Carrier sends an alert to the special channel in Slack when that request is created and we're constantly auditing its usage." |
| 18 | Mercari | "Zero Touch Production at Mercari" parts 1–2, MGLS talks (Dylan Lau) [search] | talk | 2022 | 2026-09-28 | https://www.youtube.com/watch?v=F2WEfZvQLZM | Recorded engineering account corroborating the Carrier design (request gateway, approvals, breakglass); cited as corroboration of rows 16–17, no timestamped figure taken. | Talk series titled "MGLS#72/#73 Zero Touch Production at Mercari (½, 2/2)". |
| 19 | Figma | "Designing for Security and Usability: Figma's Modern Endpoint Strategy" [search] | blog | 2025-04-08 | 2026-09-28 | https://www.figma.com/blog/figmas-modern-endpoint-strategy/ | A design-tool company runs production access as role-pre-approved JIT grants that expire in about an hour, naming usability as the reason. | Reported as: JIT role-based access via Opal lets people request access when needed; those pre-approved by role receive access that "automatically expires after a period of time", such as one hour; burdensome access procedures drive anti-patterns like shadow IT and running services from development environments. |
| 20 | Netflix | "ConsoleMe: A Central Control Plane for AWS Permissions and Access", Netflix TechBlog [search] | blog | 2021 | 2026-09-28 | https://netflixtechblog.com/consoleme-a-central-control-plane-for-aws-permissions-and-access-fd09afdd60a8 | Netflix built its own broker to manage IAM across hundreds of AWS accounts and hand out short-lived credentials through one interface. | Reported as: the Cloud Infrastructure Security Team "manages IAM permissions across hundreds of accounts"; ConsoleMe "brokers application AWS credentials to provide users with short-lived IAM credentials." |
| 21 | Netflix | ConsoleMe repository README, archive notice [direct] | source | 2026 (archive effective 2026-03-01) | 2026-09-28 | https://github.com/Netflix/consoleme | Five years after open-sourcing, Netflix archived ConsoleMe and Weep because the internal fork diverged past the point of shared maintenance. | "This repository will be archived and set to read-only on March 1, 2026"; "the open-source versions now diverge substantially from our internal implementations and no longer reflect how we use or operate these tools"; "now with 3,200+ GitHub stars." |
| 22 | GitLab | Runbooks: "Accessing the Rails Console as an SRE" [direct] | source | current | 2026-09-28 | https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/console/access.md | GitLab SREs hold standing, approval-free access to the production Rails console, the same surface its own docs call unguarded. | "Site Reliability Engineers (SREs) can access the production-rails console directly without an approval process." |
| 23 | GitLab | Runbooks: "Teleport Approver Workflow" [direct] | source | current | 2026-09-28 | https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/teleport/teleport_approval_workflow.md | For everyone below SRE the same company runs a full JIT matrix: prod read-only needs a people manager, prod read-write cannot be approved by managers at all, grants expire in 12 hours, and incidents route to on-call instead of the queue. | "Read-write access requests typically cannot be approved by engineering managers."; "Access requests are temporary and expire after 12 hours"; "If urgent access is required during a production incident, ping @sre-oncall"; approver matrix: non-prod read-only needs no approval, prod read-only needs Engineering/Security people managers. |
| 24 | GitLab | Handbook: Teleport Access policy [direct] | adr | current | 2026-09-28 | https://handbook.gitlab.com/handbook/engineering/gitlab-com/policies/teleport/ | The policy layer above the runbook: Okta-driven role baseline plus access requests, quarterly reviews, one-year immutable audit retention. | "Teleport access is managed through Okta and is provided as part of a role's baseline group assignment or through an access request with appropriate approval"; "Teleport Audit Logs must be retained for a defined period of 1 year"; "must not be modified and or deleted before the defined time." |
| 25 | GitLab | Infrastructure tracker, work item 29713: "Streamline Teleport access" [direct] | source | 2026-09-04 | 2026-09-28 | https://gitlab.com/gitlab-com/gl-infra/production-engineering/-/work_items/29713 | Three-plus years after adopting a commercial broker, the operator of a public SaaS is still openly questioning whether it should gate incident access at all. | "Teleport access processes are currently unclear, with documentation scattered across multiple locations"; proposed action: "Evaluate whether Teleport should serve as the default mechanism for granting production access to SREs and Engineers during incidents." |
| 26 | Login.gov (GSA) | Handbook: "GitLab" platform article [direct] | adr | current | 2026-09-28 | https://handbook.login.gov/articles/platform-gitlab.html | A federal identity service documents, in public, that its break-glass path exists precisely for when its own IdP is down. | "Note - If secure.login.gov is not available, existing Personal Access Tokens continue to function. We also have break-glass procedures if needed. See Runbook: GitLab Access Contingency Plan." |
| 27 | Teleport (community) | Discussion #30686: "What to do if the SSO provider is down?" [search] | source | 2023 | 2026-09-28 | https://github.com/gravitational/teleport/discussions/30686 | The vendor's own guidance for IdP failure is a standing local user held outside the SSO chain. | "If you need that break-glass access in the event of an outage, you can get the credentials and log into Teleport as this user without needing Github at all." |
| 28 | Teleport (community) | Issue #3760: "Recovery guide for Teleport HA for disaster recovery scenarios" [search] | source | 2020, standing | 2026-09-28 | https://github.com/gravitational/teleport/issues/3760 | Users have asked since 2020 for a documented recovery path for when the access broker itself is offline; discussion #48033 (2024) shows the community still building it themselves via OpenSSH trust of the Teleport CA. | Reported as: request for "guidance around break glass procedure guide for various recovery scenarios", including "if Teleport Auth cluster is offline"; see also https://github.com/gravitational/teleport/discussions/48033. |
| 29 | Kubernetes | kubeadm issue #2414 and the v1.29 super-admin.conf change [search] | adr | 2021–2023 (shipped v1.29, 2023-12) | 2026-09-28 | https://github.com/kubernetes/kubeadm/issues/2414 | The project demoted its default admin credential because a system:masters certificate bypasses RBAC and cannot be revoked short of rotating the cluster CA; the irrevocable super-credential was split into a separate file to be treated as break-glass. |
Reported as: admin.conf's credential "could not be revoked if a cluster operator lost control of it, and the only way to revoke access completely was to rotate the cluster certificate authority"; from v1.29 admin.conf carries kubeadm:cluster-admins (revocable via RBAC) while super-admin.conf keeps system:masters. |
| 30 | Rory McCune | "When is admin not admin?, when it's super-admin!" [search] | blog | 2024-01-06 | 2026-09-28 | https://raesene.github.io/blog/2024/01/06/when-is-admin-not-admin/ | Independent practitioner confirmation of row 29's mechanism and its break-glass framing. | Reported as: system:masters bypasses RBAC entirely; super-admin.conf "is a member of the system:masters group, so RBAC does not apply", making it the unrevocable fallback while admin.conf became revocable. |
| 31 | Doyensec | "Testing Zero Touch Production Platforms and Safe Proxies" [search] | blog | 2023-05-04 | 2026-09-28 | https://blog.doyensec.com/2023/05/04/testing-ztp-platforms-a-primer.html | A security consultancy that tests these platforms reports there is no mature off-the-shelf option; everyone rebuilds the proxy, and the approval flow's webhooks are a soft target. | "Many companies need secure proxy tools but are all trying to reinvent the wheel in one way or another because it's an immature market and no off-the-shelf solutions exist"; webhook risks in "the command approval flow ceremony": content spoofing, authentication weaknesses, replay. |
| 32 | Rami McCarthy | "The Path to Zero-Touch Production", fwd:cloudsec 2024 [search] | talk | 2024-06 | 2026-09-28 | https://speakerdeck.com/ramimac/the-path-to-zero-touch-production (video: https://www.youtube.com/watch?v=agzrIBY0ScQ) | Practitioner synthesis of incremental ZTP adoption for cloud-native organisations outside Google scale; frames the problem as why and how people touch prod. | Deck description: "A universal theory for incrementally moving a cloud-native org to Zero Touch Prod, with AWS production access primitives." |
| 33 | Microsoft | "Manage emergency access admin accounts", Entra ID docs [search] | vendor | current | 2026-09-28 | https://learn.microsoft.com/en-us/entra/identity/role-based-access-control/security-emergency-access | The hyperscaler's prescription for tenants: two or more break-glass accounts, at least one excluded from the conditional-access policies everything else must pass, phishing-resistant credentials, monitored sign-ins. | Reported as: create "two or more emergency access accounts"; exclude at least one from all Conditional Access policies "to prevent from getting blocked during emergencies"; current guidance prefers FIDO2 hardware passkeys or certificate-based authentication. |
| 34 | Teleport | Docs: "Configuring Break-Glass SSH Access for Disaster Recovery" [search] | vendor | current | 2026-09-28 | https://goteleport.com/docs/zero-trust-access/deploy-a-cluster/reliability/breakglass-access/ | The broker vendor documents how to survive itself: OpenSSH configured to trust the Teleport user CA keeps certificates working "even if Teleport itself is down". | Reported as: configure the OpenSSH server alongside a Teleport Agent to trust Teleport's user CA so users with valid Teleport-issued certificates can authenticate "even if Teleport itself is down." |
| 35 | Oppenheimer, Ganapathi, Patterson (UC Berkeley) | "Why Do Internet Services Fail, and What Can Be Done About It?", USITS '03 [search] | paper | 2003-03 | 2026-09-28 | https://www.usenix.org/conference/usits-03/why-do-internet-services-fail-and-what-can-be-done-about-it (PDF mirror: https://pages.cs.wisc.edu/~remzi/Classes/739/Fall2018/Papers/oppenheimer.pdf) | The two-decade-old measurement that grounds the whole field: the human operator, not the hardware, is the leading cause of service failure. | Reported as: "operator error is the largest single cause of failures in two of the three services" studied, and configuration errors were the largest category of operator errors, over 50% (nearly 100% in one service). |
| 36 | Saltzer & Schroeder (MIT) | "The Protection of Information in Computer Systems", Proc. IEEE [search] | paper | 1975 | 2026-09-28 | https://www.cs.virginia.edu/~evans/cs551/saltzer/ | The principle every access system in this guide implements or violates, stated fifty years ago. | "Every program and every user of the system should operate using the least set of privileges necessary to complete the job." |
| 37 | Govindan, Minei, Kallahalla, Koley, Vahdat (Google) | "Evolve or Die: High-Availability Design Principles Drawn from Google's Network Infrastructure", SIGCOMM 2016 [search] | paper | 2016-08 | 2026-09-28 | https://research.google/pubs/evolve-or-die-high-availability-design-principles-drawn-from-googles-network-infrastructure/ (DOI: https://dl.acm.org/doi/10.1145/2934872.2934891) | Across 100+ high-impact failure events in Google's networks, a large share happened while a management operation was in progress, the quantitative bridge from "humans touch prod" to "prod breaks". | Reported as: analysis of over 100 high-impact failure events across data centers and two WANs; "a large number of failures happen when a network management operation is in progress within the network"; ~80% of failures last between 10 and 100 minutes. |
Absences worth recording
- No rejected pull request was found in this corpus. The recorded arguments in this domain live in issues and discussions (rows 27–29), not in closed-unmerged PRs; access-control features in the open-source brokers appear to be argued before code is written, and the richest arguments (Google's, Netflix's internal fork, GitLab's 29713) happen where PRs are not public. Stated in the page as an evidence limit.
- No organisation publishes its break-glass usage rate. Mercari says it audits every use; Google says overuse is the failure mode; nobody prints the number. Named in the page as an open question.
- Segment's access-provisioning account (referenced in secondary material) could not be located as a primary source and is not cited.