Letting humans in  / field guide
Practitioner field guide · 2026-09-28

Letting humans into production

How eleven organisations grant, constrain and audit their own engineers' access to production, reconstructed from six public breach investigations (Twitter, Uber, CircleCI, Okta, Cloudflare, Microsoft) and five builder accounts (Google, Netflix, Mercari, Figma, GitLab). A reader leaves knowing the three legitimate doors into production, why approval gates protected none of the breached companies, the condition that flips every major access decision, and how to prove a break-glass path works before the day it is needed.

30 primary sources 11 organisations 6 incident reports Evidence through September 2026 Read: ~25 min
01

The territory

The problem, stated without naming a product: your own engineers must be able to act on the systems that run your business, and every mechanism that lets them act is a mechanism that mistakes and misuse act through as well.

~13%
of Google-evaluated outages could have been prevented or mitigated by Zero Touch Prod, by Google's own estimate
1,000+
Twitter employees and contractors held the internal account tools implicated in the July 2020 incident
5,000
production credentials Cloudflare rotated after four missed ones turned out to still be live
30 min
from support-file upload to detection at BeyondTrust, the customer that watched its identity plane like production

Operator access is the oldest problem in this category and still the least settled. The measurement that frames it is twenty-three years old: Oppenheimer, Ganapathi and Patterson's 2003 study of three large internet services found operator error the largest single cause of failure in two of the three, with configuration mistakes making up more than half of those errors. Google's 2016 SIGCOMM analysis of more than 100 high-impact network failures found a large share happened while a management operation was in progress. The security record points the same direction from the other side: each of the six public breach investigations in this corpus began at a person's access, not at a software vulnerability in the product itself.

The finding that reorganised this page

None of the six investigations describes an approval gate being defeated. Every report instead names a standing credential or a live session that made the gates irrelevant: a phished workforce password at Twitter, an embedded administrative credential for the privileged-access vault at Uber, a stolen single-sign-on session at CircleCI, a service-account password that had been saved into a personal browser profile at Okta, four unrotated service tokens at Cloudflare, and a signing key created in 2016 and never rotated at Microsoft. In three of the six (Uber, Okta, Microsoft) the compromised component was the access-control machinery itself. The system you build to guard production becomes the most valuable thing you own, and it ends up guarded by whatever you forgot about.

The organisations that have written about the control side converge on one rule with three clauses. Google's SREcon19 formulation, restated almost word for word by Mercari two years later, is that every production change is made by automation, prevalidated by software, or triggered through an audited break-glass mechanism. Everything else in this guide is an implementation argument about those three doors: how wide the second should be, and how loud the third must be.

Figure 1 · The three legitimate doors into production

command + session log

alarm on every use,
reviewed after

Engineer needs to
change production

Door 1
Automation makes the change,
no human session

Door 2
Tooling that pre-validates
the specific action

Door 3
Audited break-glass,
bypasses authorization

Production

Audit pipeline

command + session log

alarm on every use,
reviewed after

Engineer needs to
change production

Door 1
Automation makes the change,
no human session

Door 2
Tooling that pre-validates
the specific action

Door 3
Audited break-glass,
bypasses authorization

Production

Audit pipeline

Google's Zero Touch Prod rule, corroborated independently by Mercari: a path into production that is none of these three is an unowned risk. Sources: SREcon19 EMEA abstract, Mercari, 2022.
Diagram source

Scope, stated plainly. This guide covers interactive human access to production infrastructure: shells and consoles, cloud IAM sessions, database consoles, and the internal administrative tools that act on customer data. It deliberately does not cover product-side authorization models (policy engines deciding what end users may do), workload identity between services, or the hour after a machine credential leaks; the 2026-09-13, 2026-08-29 and 2026-09-20 guides in this collection cover those. Vendor SaaS administration appears only where the public record makes it the relevant surface.

02

How it is actually built

The common shape across Google, Netflix, Mercari, Figma and GitLab, with the divergence points marked. The emergency plane is the part most teams draw last and every incident report argues to draw first.

Figure 2 · Reference architecture for human production access

Emergency plane

Normal plane

session recording

alarm on use

Engineer on a
managed device

Identity provider
SSO + phishing-resistant MFA

Access broker
request, approval policy, TTL

Short-lived credential
certificate or token, hours

Constrained path
safe proxy or console gateway

Break-glass credential
stored outside the IdP

Production

On-call schedule and
role pre-approvals

Immutable audit log

Emergency plane

Normal plane

session recording

alarm on use

Engineer on a
managed device

Identity provider
SSO + phishing-resistant MFA

Access broker
request, approval policy, TTL

Short-lived credential
certificate or token, hours

Constrained path
safe proxy or console gateway

Break-glass credential
stored outside the IdP

Production

On-call schedule and
role pre-approvals

Immutable audit log

Reconstructed from Google's least-privilege chapter, Netflix's ConsoleMe, Mercari's Carrier, Figma's Opal deployment and GitLab's Teleport runbooks. In the normal plane every link can refuse the request; the emergency plane shares no component with it, which is the whole point.
Diagram source

The broker is the component everyone builds, buys, or regrets

Netflix built ConsoleMe to manage IAM across what its blog describes as hundreds of AWS accounts, brokering short-lived credentials from one interface. Mercari built Carrier. Figma bought Opal. GitLab bought Teleport. Doyensec, a consultancy that assesses these platforms, wrote in 2023 that "many companies need secure proxy tools but are all trying to reinvent the wheel in one way or another because it's an immature market and no off-the-shelf solutions exist". Three years on, that judgement still describes the corpus: every organisation here runs a different broker, and two of the five built their own.

Sources: Netflix, Mercari, Figma, Doyensec

What the operator touches diverges most

Google's safe proxies review and run individual commands, so the unit of access is one vetted action. GitLab's SREs get a full Ruby console on the production application, which the runbook grants "directly without an approval process". Mercari's Carrier sits between, granting scoped permissions on request. This is the widest divergence in the corpus, and it is a bet about where risk lives: Google bets the risky object is the command, GitLab bets it is the delay during an incident.

Sources: Google Cloud docs, GitLab runbooks

The audit pipeline is a retention policy, not a log line

GitLab's Teleport policy is explicit in a way most handbooks are not: audit logs are retained one year, "must not be modified and or deleted" before that, and access to the audit data is itself least-privilege. Google's chapter argues the cultural half: "Without cultural reinforcement, audits can become rubber stamps, and breakglass use can become an everyday occurrence, losing its sense of importance or urgency." An audit trail nobody reads quietly converts the third door into an unlocked one.

Sources: GitLab handbook, Google, BSRS ch. 5

The emergency plane is standard practice, not an exception

Microsoft's Entra guidance prescribes two or more emergency-access accounts with at least one excluded from the Conditional Access policies everything else must satisfy. Teleport's docs describe configuring OpenSSH to trust the Teleport CA so access still works even if Teleport itself is down. Login.gov, a federal identity service, documents publicly that its GitLab contingency runbook exists for the day the IdP is unavailable. Kubernetes redesigned kubeadm in v1.29 so the unrevocable system:masters credential lives in a separate super-admin.conf, to be treated exactly like a break-glass account. Four independent designs, one conclusion: the fallback must not depend on anything it is a fallback for.

Sources: Microsoft, Teleport docs, Login.gov, kubeadm #2414

Why the emergency plane earns a whole plane rather than a checkbox: the normal plane is circular. The broker fronts production, the identity provider gates the broker, and the identity provider is itself a production system, run by you or by a vendor whose own support tooling sits inside your trust boundary. The October 2023 Okta incident made the vendor half of that circle concrete: files uploaded to the identity vendor's support system exposed customers' session material, and Cloudflare's February 2024 report traces its own intrusion to credentials exposed in that same event. When the circle fails, the only path that still works is the one that was never inside it.

Figure 3 · The circular dependency the break-glass path exists to break

gates

fronts

runs or hosts

operate

the only edge that
survives the circle

Identity provider

Access broker

Production

Vendor's own support
and admin systems

Break-glass credential
offline, alarmed, rehearsed

gates

fronts

runs or hosts

operate

the only edge that
survives the circle

Identity provider

Access broker

Production

Vendor's own support
and admin systems

Break-glass credential
offline, alarmed, rehearsed

Access to production depends on systems that are themselves production; the emergency credential is useful precisely because it sits outside the circle. Sources: Login.gov handbook, Teleport discussion #30686, Okta root-cause report.
Diagram source

One more architectural fact, reported rather than inferred: this layer has a short half-life. Netflix open-sourced ConsoleMe and Weep in March 2021; the repository's README now announces archival on March 1, 2026, because "the open-source versions now diverge substantially from our internal implementations and no longer reflect how we use or operate these tools". GitLab adopted Teleport years ago and opened a tracker item in September 2026 proposing to "evaluate whether Teleport should serve as the default mechanism for granting production access to SREs and Engineers during incidents". An access broker is not a purchase; it is an ongoing negotiation between the security team and everyone else, and the corpus shows that negotiation reopening every few years.

03

The decisions that matter

Five forks in the road, with who went which way, the stated reason, and the condition that flips the answer. The disagreements here are the best material in the corpus, because both sides are competent and both wrote down why.

Do engineers hold standing access, or request it just in time?

Chosen: JIT with expiry
  • Figma: role-pre-approved requests through Opal, expiring on the order of an hour (2025)
  • GitLab, below SRE level: Teleport requests that expire after 12 hours (current runbook)
  • Netflix: short-lived IAM credentials brokered per request (2021)
Rejected: broad standing roles
  • The Twitter report is the canonical exhibit: over 1,000 people held the account tools, and the report ties the incident's reach directly to that population
  • Figma names the second cost: burdensome access drives shadow IT and production work from unmanaged environments
Flips when
  • Response latency is the product: GitLab leaves SREs standing console access because minutes matter mid-incident, and routes urgent grants to on-call instead of a queue
  • The honest form of the flip is a narrow, named, monitored population, never a wide default

Does the emergency path live inside or outside the SSO chain?

Chosen: outside, alarmed
  • Microsoft Entra: two or more accounts, at least one excluded from Conditional Access, hardware-key credentials, monitored sign-ins
  • Teleport guidance: a local user whose credentials work "without needing Github at all" when SSO is down
  • Login.gov: a public contingency runbook for the day the IdP is unavailable
Rejected: fallback via the IdP
  • Structurally circular: the failure being planned for removes the path that was planned
  • kubeadm's pre-1.29 design shows the other extreme, an unrevocable super-credential used as the daily driver
Flips when
  • It does not flip; what varies is the credential's blast radius. Kubernetes' answer: keep the unrevocable credential but demote it to a sealed file never used routinely

In an emergency, does anyone approve, or does the alarm replace approval?

Chosen: bypass now, audit loudly
  • Mercari's Carrier: "BreakGlass is a way for developers to get permissions without reviewer approval in emergency situations", with an alert to a Slack channel on every use and constant auditing
  • Google's book defines breakglass as bypassing the authorization system completely
Rejected: unilateral access to everything
  • Google's production-protection page states that unilateral access to foundational services is not allowed, and that even emergency access requires approval from other personnel
  • The same book flags the known weakness of multi-party approval: the approving parties usually work from the same managed fleet, so it is not independent when that fleet is the thing in question
Flips when
  • Blast radius decides it. Foundational services whose failure reaches everyone keep approval even in emergencies; narrower surfaces get the alarmed bypass. GitLab encodes this at fine grain: managers may approve a customer-infrastructure write only when it affects a single customer, no bulk changes

Figure 4 · A decision tree the corpus supports

yes

no

yes

no

no

yes

yes

no

Can automation or existing
tooling make this change?

Use it. No human
session is created

Read-only?

JIT read grant,
auto-expiry, light approval

Active incident?

JIT write grant: approval by a
practitioner, not a manager, plus TTL

IdP and broker
still reachable?

Pre-approved on-call role;
approval rides the page, not a queue

Break-glass credential.
Alarm fires now, review follows

yes

no

yes

no

no

yes

yes

no

Can automation or existing
tooling make this change?

Use it. No human
session is created

Read-only?

JIT read grant,
auto-expiry, light approval

Active incident?

JIT write grant: approval by a
practitioner, not a manager, plus TTL

IdP and broker
still reachable?

Pre-approved on-call role;
approval rides the page, not a queue

Break-glass credential.
Alarm fires now, review follows

Terminal nodes are actions. The tree encodes GitLab's approval matrix and expiry, Mercari's alarmed break-glass, and Google's automation-first rule; your numbers will differ, the ordering should not. Sources: GitLab, Mercari, Google.
Diagram source
DecisionChosenRejectedBecauseFlips whenEvidence
Standing vs JITJIT, expiring grants (1 h Figma, 12 h GitLab)Broad standing rolesThe holder population is the exposed surface; 1,000+ at TwitterNamed on-call responders during incidentsFigma 2025, NYDFS 2020
Who approves prod writesPractitioners (SRE/DBRE)Engineering managersGitLab: managers may approve prod read-only, but "read-write access requests typically cannot be approved by engineering managers"Single-customer scope (the CustomersDot carve-out)GitLab runbook
Emergency path locationOutside the SSO chain, alarmedFallback through the same IdPThe dependency is circular; the IdP is production tooDoes not flip; only the credential's scope variesMicrosoft, Teleport #30686
Emergency approvalAlarmed bypass for narrow surfacesUnilateral access to foundational servicesSpeed against blast radiusRadius reaches everyone: approval survives even emergencies (Google)Mercari 2022, Google docs
Build vs buy the brokerBoth, in roughly equal numbersAssuming the market is settledDoyensec 2023: immature market, "no off-the-shelf solutions exist"Willingness to own the fork forever: Netflix archived its open-source broker in 2026 once internal and public versions divergedDoyensec, ConsoleMe README
Unit of accessCommand via safe proxy (Google)Raw shell as the defaultA reviewed command has a bounded blast radius; a full console has noneIncident latency dominates and the operator population is small and senior (GitLab's console)Google, GitLab
Credential revocabilityRevocable group binding (kubeadm:cluster-admins)Daily use of system:mastersThe old admin.conf could only be revoked by rotating the cluster CANever; the unrevocable form survives only as sealed break-glasskubeadm #2414, McCune 2024

Where the sources genuinely disagree: Google against GitLab, on standing operator access. Google removed it and measured the reliability dividend. GitLab keeps a console one SSH command away from its senior operators, publishes that fact, and is still arguing with itself about it in the open; the September 2026 tracker item is the argument's latest round. The context difference that explains both positions: Google amortises tooling across thousands of operators and can afford to wrap nearly every operational action in an API; GitLab runs one very deep application whose incidents are often only diagnosable in the application's own console. If your operational surface looks like a thousand services, copy Google. If it looks like one deep application, GitLab's compromise (standing access for a small senior group, JIT for everyone else, one year of immutable audit) is the honest one, and its cost is that you inherit the argument too.

04

What broke in production

Six published investigations, grouped into the four failure classes they actually form. The classes matter more than the individual incidents: each names a place where an approval gate provides no protection at all.

Read the six reports together and four patterns repeat. Class 1: the help desk is an authentication surface. Attackers do not request access; they talk to someone who already has it. Class 2: a standing credential makes every gate downstream of it irrelevant. The forgotten service account, the embedded script password, the unrotated signing key. Class 3: the session outlives the check. Multi-factor authentication happens at login; the artefacts of a logged-in session persist afterwards. Class 4: the access vendor is inside the trust boundary. The support plane of the company that runs your identity is, in effect, part of your production.

Class 1 · Postmortem

Twitter, July 2020: a phone call reached the account tools

AssumptionInternal administrative tools were safe because only staff could reach them, behind a VPN and multi-factor login.
What happenedPer the regulator's report, callers posing as the IT help desk, against the backdrop of genuine VPN problems during the remote-work transition, led employees to a lookalike login page. Over 1,000 employees and contractors held the account tools, which widened the set of people worth targeting.
Blast radiusTakeover of high-profile accounts, at least $118,000 taken through the scam, and a formal state investigation. Twitter had no CISO at the time.
FixThe report records that Twitter reduced the tool-holding population immediately afterward, accepting slower support workflows as the cost.
Design ruleThe number of people holding a privileged tool is itself the attack surface. A credential-entry prompt that arrives by phone should be treated as hostile, and only phishing-resistant factors survive that assumption.
Class 2 · Postmortem

Uber, September 2022: the privilege system held a key to itself

AssumptionMulti-factor login plus a privileged-access-management vault meant no single discovered credential could open the wider estate.
What happenedPublic reporting describes repeated login prompts sent to a contractor, followed by a message impersonating Uber IT, ending in one approval. Once on the internal network, the attacker reportedly found an administrative credential for the privileged-access tool stored in a script on an internal share, which then unlocked further systems.
Blast radiusAccess to multiple internal tools and cloud consoles per reporting; the analysis notes Uber's public statement of no evidence of access to sensitive user data such as trip history.
FixPublic reporting focuses on the root cause rather than a durable fix; the transferable lesson is architectural.
Design ruleA privileged-access system that can be unlocked by a static secret stored elsewhere provides the illusion of a vault, not a vault. The credentials to the access system belong under stricter control than the systems it fronts, and static admin secrets in scripts are the specific anti-pattern.
Class 3 · Postmortem

CircleCI, January 2023: the session outlived the login check

AssumptionAn engineer authenticating with multi-factor SSO to a role that could mint production tokens was adequately protected by that login step.
What happenedPer CircleCI's report, malware on one engineer's laptop, undetected by antivirus, captured an already-authenticated SSO session. Because the session existed after the multi-factor step, the login control was not in the path; the targeted engineer's elevated permissions then reached a subset of production data stores.
Blast radiusExposure of customer secrets held in the platform; roughly 19 days from device compromise (Dec 16) to a full understanding of scope (Jan 4), with the first external signal coming from a customer noticing unusual activity.
FixShortened session lifetimes, tightened detection on the paths that mint production credentials, and a broad rotation of customer secrets.
Design ruleMulti-factor authentication protects the moment of login, not the session that follows it. For any role that can issue production credentials, the session lifetime and device posture are the real controls, and detection has to watch the credential-minting path, not only the door.
Class 4 · Postmortem

Okta, October 2023: the identity vendor's support plane

AssumptionA service account confined to the customer-support case system was low-risk because of where it lived.
What happenedOkta's root-cause report states that the service account's credentials had been saved into an employee's personal profile in a browser on a managed laptop; the most likely exposure was the compromise of that personal account or device. Support files uploaded by customers contained session material usable against those customers' own tenants.
Blast radiusFiles for 134 customers (under 1%) exposed across roughly three weeks. BeyondTrust, watching its own identity plane, flagged the resulting activity within 30 minutes; Okta confirmed the breach about two weeks after the first report.
FixOkta disabled the account and applied a browser configuration preventing sign-in to a personal profile on managed laptops; the detection-time gap between vendor and customer is the durable lesson.
Design ruleYour identity vendor's support systems are inside your trust boundary. Treat material shared with any vendor support channel as if it will be exposed, and monitor your own identity plane for anomalies rather than waiting for the vendor to tell you.
Class 2 · Postmortem

Cloudflare, November 2023: four credentials nobody rotated

AssumptionCredentials exposed in an upstream vendor breach had all been rotated; the few skipped ones were believed unused.
What happenedPer Cloudflare's report, one service token and three service accounts, out of thousands, were missed during rotation because they were mistakenly believed unused. They were still live and admitted an actor to the internal Atlassian estate (Jira and Confluence).
Blast radiusAccess to internal wiki and issue data; Cloudflare reports no customer data or systems were affected. Remediation rotated about 5,000 production credentials and triaged 4,893 systems.
FixCompany-wide credential rotation, reimaging, and closing the inventory gap that let credentials be labelled "unused" without proof.
Design rule"Believed unused" is not a security state. A credential you cannot prove is dead is alive, and rotation is only as complete as the inventory it runs against; the cost of losing inventory certainty is a full reset.
Class 2 · Postmortem

Microsoft, 2023: a signing key from 2016

AssumptionA consumer signing key could remain in service indefinitely; automated rotation for it was not prioritised.
What happenedPer the Cyber Safety Review Board, tokens signed by a key Microsoft created in 2016 and never rotated were used to reach Exchange Online mailboxes. The Board found Microsoft could not establish how the key left its custody, and described a "cascade of avoidable errors".
Blast radiusMailboxes of 22 organisations and over 500 individuals, including senior US officials. The Board concluded Microsoft's security culture "was inadequate and requires an overhaul".
FixThe Board's central recommendation was to prioritise automated key rotation, the control whose absence let a nine-year-old key still sign valid tokens.
Design ruleA credential's danger is its age times its scope. A signing key is the most privileged credential you hold; if you cannot rotate it automatically, you cannot claim to control it, and "we don't know how it was taken" is the likely end state.
What the pattern says to design for

The six reports name assumptions, not exploits, so the transferable content is a checklist of beliefs to stop holding: that internal tools are safe because they are internal; that a vault protects you when its own key is stored in the clear; that multi-factor login protects a session; that "unused" means dead; that a key can live forever; and that a vendor's support desk is outside your perimeter. An approval workflow addresses none of these. Zero standing privilege, short sessions, automated rotation and credential inventory do.

05

Numbers you can plan against

The quantities the corpus actually publishes are population sizes, grant lifetimes, detection times and remediation counts, not latency or throughput. Each row names whether it is measured, claimed or an estimate, with its date.

QuantityValueWhereKindAs ofSource
Outages preventable/mitigable by Zero Touch Prod~13%GoogleClaimed (vendor)2026Google Cloud docs
People holding the internal account tools1,000+TwitterMeasured (regulator)2020NYDFS
JIT grant lifetime, per request~1 hFigma (Opal)Reported2025Figma
JIT grant lifetime, per request12 hGitLab (Teleport)Reported (runbook)2026GitLab
Audit-log retention, immutable1 yrGitLabReported (policy)2026GitLab handbook
AWS accounts under one brokerhundredsNetflixReported2021Netflix
Detection time, customer-side30 minBeyondTrust (Okta event)Reported2023BeyondTrust
Dwell time, device compromise to scope understood~19 dCircleCIDerived from report timeline (Dec 16 to Jan 4)2023CircleCI
Credentials rotated after the incident~5,000CloudflareReported2023Cloudflare
Systems triaged after the incident4,893CloudflareReported2023Cloudflare
Unrotated credentials that caused it4CloudflareReported2023Cloudflare
Age of the unrotated signing key~7 yrMicrosoftDerived (2016 to 2023 use)2023CSRB
Mailboxes reached via that key22 orgs / 500+ peopleMicrosoftMeasured (review board)2023CSRB
Operator error as share of service failureslargest of 3 causes in 2 of 3 servicesBerkeley studyMeasured (paper)2003Oppenheimer et al.
Emergency-access accounts recommended≥ 2Microsoft EntraReported (guidance)2026Microsoft
Read these carefully

The 13% figure is Google's own, a vendor claim with no independent replication; treat it as directional. The grant lifetimes (1 h, 12 h) are two organisations' current settings, not an industry norm, and both are configuration values that will drift; check the live runbook before quoting GitLab's. The break-glass usage rate, the number that would tell you whether the third door is being overused, is published by nobody in this corpus; it is the most important number that does not exist, and it is section 4's open question turned into an operational metric you should measure for yourself.

06

The evidence wall

Every source behind this page, graded. Filter by kind.

A note on how these were read: this session's network reached GitHub, GitLab, raw source hosts and cloud.google.com directly; the other hosts were confirmed through server-side web search that returns live page excerpts. Every URL was located in this session, none cited from memory. The full ledger, one row per claim with the supporting quote, ships beside this page as sources.md.

Case study Google2020

Building Secure and Reliable Systems, ch. 5 — Design for Least Privilege

Defines Zero Touch interfaces as removing direct human access to reduce outages, defines break-glass as a complete authorization bypass, and names the weakness of multi-party approval: the parties usually share one managed fleet.

Carry forwardHuman access is a reliability problem as much as a security one; design the audit culture, not just the log.
google.github.io/building-secure-and-reliable-systems/raw/ch05.html
Talk Google / USENIX2019-10

Zero Touch Prod: Towards Safer and More Secure Production Environments

The canonical rule: every production change is made by automation, prevalidated by software, or triggered through an audited break-glass mechanism. Three doors, and only three.

Carry forwardUse the three-door test to audit your own access paths; anything that fits none is unowned risk.
usenix.org/conference/srecon19emea/presentation/czapinski
Vendor Google Cloud2026

How Google protects its production services

Estimates ~13% of Google-evaluated outages preventable or mitigable by Zero Touch Prod, and states that unilateral access to foundational services is not allowed, even in emergencies.

Carry forwardTie blast radius to approval: the wider the reach, the more emergency access still needs a second person.
docs.cloud.google.com/docs/security/production-services-protection
Postmortem CircleCI2023-01

CircleCI incident report for January 4, 2023

Malware on an engineer's laptop captured a live, already-authenticated SSO session, putting the login-time multi-factor check out of the path; the targeted role's permissions reached production data.

Carry forwardProtect the session and the device for any role that can mint production credentials, not just the login.
circleci.com/blog/jan-4-2023-incident-report
Postmortem Okta2023-11

Unauthorized Access to Okta's Support Case Management System

A support-system service account's credentials were saved into an employee's personal browser profile on a managed laptop; customer-uploaded support files held reusable session material.

Carry forwardThe identity vendor's support plane is inside your trust boundary; monitor your own tenant rather than waiting to be told.
sec.okta.com/articles/2023/11/unauthorized-access-oktas-support-case-management-system-root-cause
Postmortem Cloudflare2024-02

Thanksgiving 2023 security incident

Four credentials out of thousands were missed in a rotation because they were believed unused; they were live. Remediation rotated ~5,000 credentials and triaged 4,893 systems.

Carry forward"Believed unused" is not dead; rotation is only as complete as the inventory behind it.
blog.cloudflare.com/thanksgiving-2023-security-incident
Postmortem CISA / CSRB2024-03

Review of the Summer 2023 Microsoft Exchange Online Intrusion

Tokens signed by a 2016 key that was never rotated reached mailboxes of 22 organisations and 500+ people; the board could not establish how the key was taken and called the security culture inadequate.

Carry forwardDanger equals age times scope; if you cannot rotate a signing key automatically, you do not control it.
cisa.gov/…/CSRBReviewOfTheSummer2023MEOIntrusion508.pdf
Eng blog GitGuardian2022

Uber Breach 2022 — Everything You Need to Know

Reconstructs the September 2022 incident: after MFA-fatigue plus a social-engineering call, an administrative credential for the privileged-access tool, stored in a script on an internal share, opened downstream systems.

Carry forwardGuard the credentials to the access system harder than the systems it fronts; no static admin secrets in scripts.
blog.gitguardian.com/uber-breach-2022
Eng blog Netflix2021

ConsoleMe: A Central Control Plane for AWS Permissions and Access

One broker across hundreds of AWS accounts, handing out short-lived credentials through a self-service interface. The build case for centralising the access plane.

Carry forwardShort-lived credentials brokered centrally beat long-lived keys scattered per account.
netflixtechblog.com/consoleme-a-central-control-plane…
Source Netflix2026

Netflix/consoleme — repository README, archive notice

Archived March 1, 2026 because the open-source version "diverge[d] substantially from our internal implementations". 3,200+ stars. The half-life of an access broker, in the maintainer's own words.

Carry forwardAdopting an access broker means owning a fork that drifts; budget for the divergence.
github.com/Netflix/consoleme
Eng blog Mercari2022-01

Shifting to Zero Touch Production

An organisation outside Google independently adopts the same triad and states the rule for manual changes: "an approval or audited break glass system should be used".

Carry forwardThe three-door rule is not Google-specific; a mid-size org can adopt it verbatim.
engineering.mercari.com/en/blog/entry/20220126-shifting-to-zero-touch-production
Eng blog Mercari2022-02

Promote Zero Touch Production — further features of Carrier

Carrier's break-glass grants permissions without reviewer approval in emergencies, but alarms a dedicated Slack channel on every use and audits it continuously.

Carry forwardMake the emergency door approval-free but loud; the alarm, not the gate, is the control.
engineering.mercari.com/en/blog/entry/20220201-promote-zero-touch-production…
Eng blog Figma2025-04

Designing for Security and Usability: Figma's Modern Endpoint Strategy

Role-pre-approved just-in-time access via Opal, auto-expiring around an hour, framed against the anti-patterns burdensome access creates: shadow IT and prod work from unmanaged environments.

Carry forwardUsability is a security control here; access too painful to get is routed around.
figma.com/blog/figmas-modern-endpoint-strategy
Source GitLab2026

Runbooks — Teleport Approver Workflow & Rails Console access

A full JIT matrix in the open: managers may approve prod read-only, not read-write; grants expire after 12 hours; incidents route to on-call. Separately, SREs get the Rails console "directly without an approval process".

Carry forwardA defensible split: standing access for a small senior group, JIT for everyone else.
gitlab.com/gitlab-com/runbooks/…/teleport_approval_workflow.md
Policy GitLab2026

Handbook — Teleport Access policy

Okta-driven role baseline plus access requests, quarterly access reviews, and audit logs retained one year, immutable, with least-privilege access to the audit data itself.

Carry forwardImmutable, time-bounded audit retention is the difference between a log and an audit trail.
handbook.gitlab.com/handbook/engineering/gitlab-com/policies/teleport
Source GitLab2026-09

Infra tracker #29713 — Streamline Teleport access

Years after adopting a commercial broker, the operator openly asks whether Teleport "should serve as the default mechanism for granting production access… during incidents". The negotiation reopening in public.

Carry forwardExpect to revisit the access model every few years; it is never finished.
gitlab.com/gitlab-com/gl-infra/production-engineering/-/work_items/29713
Design record Kubernetes2023-12

kubeadm #2414 & the v1.29 super-admin.conf split

The default admin credential was demoted because a system:masters certificate bypasses RBAC and can only be revoked by rotating the cluster CA; the unrevocable form was moved to a separate file to be treated as break-glass.

Carry forwardKeep the unrevocable super-credential sealed and unused; make the daily driver revocable.
github.com/kubernetes/kubeadm/issues/2414
Source Teleport (community)2023

Discussion #30686 — What to do if the SSO provider is down?

The vendor's guidance for IdP failure is a standing local user that logs in "without needing Github at all"; issue #3760 shows users asking for a documented recovery path since 2020.

Carry forwardThe emergency account must authenticate through a path the IdP outage cannot touch.
github.com/gravitational/teleport/discussions/30686
Vendor Microsoft2026

Entra ID — Manage emergency access admin accounts

Two or more emergency accounts, at least one excluded from all Conditional Access policies, phishing-resistant credentials, monitored sign-ins. The mainstream prescription for the emergency plane.

Carry forwardTwo break-glass accounts, not one; monitor their use as the highest-signal alert you own.
learn.microsoft.com/en-us/entra/identity/role-based-access-control/security-emergency-access
Practitioner BeyondTrust2023-10

BeyondTrust Discovers Breach of Okta Support Unit

The customer that treated its identity plane like production flagged the anomalous session activity within 30 minutes of the support-file upload, roughly two weeks before the vendor confirmed the breach.

Carry forwardDetection you own beats disclosure you wait for; instrument the identity plane yourself.
beyondtrust.com/blog/entry/okta-support-unit-breach
Practitioner Doyensec2023-05

Testing Zero Touch Production Platforms and Safe Proxies

A consultancy that assesses these platforms reports no mature off-the-shelf option exists and flags the approval flow's webhooks (spoofing, weak auth, replay) as the soft target.

Carry forwardThe approval ceremony's webhooks are attack surface; authenticate and anti-replay them.
blog.doyensec.com/2023/05/04/testing-ztp-platforms-a-primer.html
Talk R. McCarthy, fwd:cloudsec2024-06

The Path to Zero-Touch Production

A synthesis of incrementally moving a cloud-native org to Zero Touch Prod, framed around why and how people touch production and what to do about it, with AWS access primitives.

Carry forwardZTP is reachable incrementally; you do not need Google's scale to start.
speakerdeck.com/ramimac/the-path-to-zero-touch-production
Paper Oppenheimer et al., USITS2003

Why Do Internet Services Fail, and What Can Be Done About It?

Operator error was the largest single failure cause in two of three large services studied, with configuration errors the largest operator-error category. The empirical floor under the whole topic.

Carry forwardReducing human touch is a reliability lever with two decades of evidence, not only a security one.
usenix.org/conference/usits-03/why-do-internet-services-fail-and-what-can-be-done-about-it
Paper Google, SIGCOMM2016

Evolve or Die: High-Availability Design Principles from Google's Network

Across 100+ high-impact failures, a large share occurred while a management operation was in progress; ~80% lasted between 10 and 100 minutes. Quantifies the change-path risk.

Carry forwardThe change path, not the steady state, is where operator-driven failure concentrates.
research.google/pubs/evolve-or-die…
Paper Saltzer & Schroeder1975

The Protection of Information in Computer Systems

The origin of least privilege: "Every program and every user of the system should operate using the least set of privileges necessary to complete the job." Every design in this guide implements or violates it.

Carry forwardLeast privilege is fifty years old; the novelty is only in making it low-friction enough to keep.
cs.virginia.edu/~evans/cs551/saltzer
Practitioner R. McCune2024-01

When is admin not admin?, when it's super-admin!

Independent confirmation of the kubeadm change: system:masters bypasses RBAC, so super-admin.conf is the unrevocable fallback while admin.conf became revocable.

Carry forwardKnow which of your credentials bypass your own authorization layer; those are break-glass, not daily tools.
raesene.github.io/blog/2024/01/06/when-is-admin-not-admin
Evidence limits

No rejected pull request appears in this corpus: the recorded arguments in this domain live in issues, discussions and internal design (Google's book, Netflix's fork, GitLab's tracker), not in closed-unmerged PRs, and the richest of those are not public. No organisation publishes its break-glass usage rate. Both gaps are stated where they bear on a claim, and neither is padded over.

07

Build a miniature, then productionise it

Seven rungs, from an afternoon's script to a production-shaped access plane. The line from toy to real is crossed at rung 4, where the emergency path and the audit trail appear; everything after that is the operational surface the incidents say matters most.

Inventory who can reach production today

List every standing path into one production system: human accounts, service accounts, long-lived keys, shared credentials, and the tools that front them. Do it for one system, by hand.

Done when: you have a single list and at least one entry surprises you.  Teaches: that Cloudflare's "believed unused" problem is already yours; you cannot rotate what you have not enumerated.

Broker a short-lived credential

Write a small service that, on an authenticated request, issues a credential that expires in an hour (an STS token, a short-TTL certificate, whatever your platform offers) instead of handing out a standing key.

Done when: the credential stops working after its TTL, observably.  Teaches: the core of ConsoleMe and Teleport; expiry is the property that makes a stolen credential a smaller problem.

Put an approval and a TTL on a write

Add a request/approve step in front of a write-capable grant, with the grant auto-expiring. Encode GitLab's rule: a different person approves, and read-only needs a lighter approver than read-write.

Done when: an unapproved request cannot act, and an approved one stops acting after the TTL.  Teaches: just-in-time access, and why the approver's identity is part of the policy.

Build the emergency plane, separately — the toy-to-real line

Create a break-glass credential that authenticates through a path your normal IdP outage cannot touch, stored offline. Wire it so that any use fires a loud alarm (a dedicated channel, a page) before anyone reviews it.

Done when: you can log in with it while the IdP is deliberately switched off, and using it pages you.  Teaches: the circular-dependency lesson (Fig. 3); the fallback must share nothing with what it backs up.

Make the audit trail immutable and time-bounded

Record every grant, every command or session, and every break-glass use to a store the operators cannot edit or delete, retained for a fixed period. Copy GitLab's policy: one year, no deletion, least-privilege access to the audit data itself.

Done when: an operator with full production access still cannot alter their own audit record.  Teaches: the difference between a log and an audit trail, and why Google says the trail is a cultural artefact.

Rehearse the failure paths on purpose

Run a game day: switch off the IdP and recover using only the break-glass path; then revoke a compromised operator's access and confirm every standing credential of theirs is dead. Time both.

Done when: both drills succeed without improvisation, and you have the two times written down.  Teaches: that a break-glass path you have never used is a hypothesis; the Twitter and Cloudflare fixes were both revocation exercises no one had rehearsed.

Instrument the numbers nobody publishes

Emit metrics for break-glass usage rate, time-to-grant during incidents, and the size of the standing-access population. Alert when any of them drifts the wrong way.

Done when: a rising break-glass rate or a growing standing population shows up on a dashboard before it shows up in an incident.  Teaches: that the control degrades silently (Google's "everyday occurrence"), and the only defence is measuring the thing the corpus refuses to publish.

08

Keep hunting

The queries that actually surfaced this material, grouped by what they turn up. The page will go stale; the method does not. The strongest move in this domain is to read the incident report and the builder's own runbook, never the summary of either.

Incident reports (the highest-value tier)

  • <company> incident report "root cause" access session
  • <vendor> "we failed to rotate" OR "believed unused" credentials
  • "support system" breach service account "personal" root cause
  • CSRB OR NYDFS report <company> filetype:pdf

How builders actually run it

  • <company> engineering "zero touch production" OR "break glass"
  • <company> "just-in-time" access "expires" production
  • site:handbook.<company>.com teleport OR "access request" approval
  • "we built" OR "we replaced" access broker short-lived credentials

The design arguments (issues, RFCs, runbooks)

  • repo:<org>/<repo> is:issue "break glass" OR "SSO is down"
  • <project> "super-admin" OR system:masters revoke RBAC
  • path:docs runbook "when IDP is down" OR contingency access
  • "emergency access" accounts exclude conditional access break glass

The evidence base under it all

  • operator error largest cause internet service failures
  • failures during "management operation" network Google SIGCOMM
  • "least privilege" Saltzer Schroeder 1975 definition
  • testing "zero touch production" safe proxies webhook
09

References

  1. Barrett, Joyner & Ward, "Design for Least Privilege", Building Secure and Reliable Systems, ch. 5 O'Reilly / Google, 2020. Checked 2026-09-28.
  2. Czapiński & Wolafka, "Zero Touch Prod: Towards Safer and More Secure Production Environments" USENIX SREcon19 EMEA, October 2019. Checked 2026-09-28.
  3. "How Google protects its production services" Google Cloud documentation, current. Checked 2026-09-28.
  4. "Twitter Investigation Report" New York State Department of Financial Services, 2020-10-14. Checked 2026-09-28.
  5. "CircleCI incident report for January 4, 2023 security incident" CircleCI, 2023-01-13. Checked 2026-09-28.
  6. Bradbury, "Unauthorized Access to Okta's Support Case Management System: Root Cause and Remediation" Okta Security, 2023-11-03. Checked 2026-09-28.
  7. "Thanksgiving 2023 security incident" Cloudflare, 2024-02-01. Checked 2026-09-28.
  8. "Review of the Summer 2023 Microsoft Exchange Online Intrusion" Cyber Safety Review Board / CISA, 2024-03-20. Checked 2026-09-28.
  9. "Uber Breach 2022 – Everything You Need to Know" GitGuardian, 2022. Checked 2026-09-28.
  10. Castrapel & Dubey, "ConsoleMe: A Central Control Plane for AWS Permissions and Access" Netflix Technology Blog, 2021. Checked 2026-09-28.
  11. Netflix/consoleme — repository and archive notice Netflix / GitHub, archived 2026-03-01. Checked 2026-09-28.
  12. "Shifting to Zero Touch Production" Mercari Engineering, 2022-01-26. Checked 2026-09-28.
  13. "Promote Zero Touch Production – further features of Carrier" Mercari Engineering, 2022-02-01. Checked 2026-09-28.
  14. "Designing for Security and Usability: Figma's Modern Endpoint Strategy" Figma, 2025-04-08. Checked 2026-09-28.
  15. "Teleport Approver Workflow" runbook GitLab Runbooks, current. Checked 2026-09-28.
  16. "Accessing the Rails Console as an SRE" runbook GitLab Runbooks, current. Checked 2026-09-28.
  17. "Teleport Access policy" GitLab Handbook, current. Checked 2026-09-28.
  18. "Streamline Teleport access" (work item 29713) GitLab Infrastructure tracker, 2026-09-04. Checked 2026-09-28.
  19. kubeadm issue #2414 (super-admin.conf / admin.conf split) Kubernetes / GitHub, shipped v1.29 (2023-12). Checked 2026-09-28.
  20. McCune, "When is admin not admin?, when it's super-admin!" raesene.github.io, 2024-01-06. Checked 2026-09-28.
  21. Teleport discussion #30686, "What to do if the SSO provider is down?" (and #3760, #48033) Gravitational / GitHub, 2023. Checked 2026-09-28.
  22. "Configuring Break-Glass SSH Access for Disaster Recovery" Teleport documentation, current. Checked 2026-09-28.
  23. "Manage emergency access admin accounts" Microsoft Entra ID documentation, current. Checked 2026-09-28.
  24. "GitLab" platform article (break-glass / IdP-down runbook) Login.gov Handbook (GSA), current. Checked 2026-09-28.
  25. "BeyondTrust Discovers Breach of Okta Support Unit" BeyondTrust, 2023-10-20. Checked 2026-09-28.
  26. "Testing Zero Touch Production Platforms and Safe Proxies" Doyensec, 2023-05-04. Checked 2026-09-28.
  27. McCarthy, "The Path to Zero-Touch Production" fwd:cloudsec, June 2024 (video). Checked 2026-09-28.
  28. Oppenheimer, Ganapathi & Patterson, "Why Do Internet Services Fail, and What Can Be Done About It?" USENIX USITS '03, March 2003. Checked 2026-09-28.
  29. Govindan et al., "Evolve or Die: High-Availability Design Principles Drawn from Google's Network Infrastructure" ACM SIGCOMM 2016. Checked 2026-09-28.
  30. Saltzer & Schroeder, "The Protection of Information in Computer Systems" Proceedings of the IEEE, 1975. Checked 2026-09-28.