Letting humans into production
How organisations grant, constrain and audit their own engineers' interactive access to production, reconstructed from six public breach investigations (Twitter, Uber, CircleCI, Okta, Cloudflare, Microsoft) and five builder accounts (Google, Netflix, Mercari, Figma, GitLab).
A field guide to the human-access plane: the three legitimate doors into production (automation, pre-validated tooling, audited break-glass), why none of the six investigated breaches defeated an approval gate, and the conditions that flip every major access decision. A reader leaves with a decision tree for standing versus just-in-time access, a four-class failure catalogue with transferable design rules, a seven-rung build ladder that crosses from toy to production-shaped at the emergency plane, and the search method that surfaced the evidence.
Not one of the six investigated breaches describes an approval gate being defeated; every one walked around the gates on a standing credential or a live session, and in three of the six the compromised component was the access-control machinery itself.
What you get out of it
- Approval workflows protected none of the breached companies; standing credentials, stolen sessions and unrotated keys made the gates irrelevant, so zero standing privilege and short sessions matter more than the approval step.
- The break-glass path must share nothing with the systems it backs up: four independent designs (Microsoft Entra, Teleport, Login.gov, Kubernetes kubeadm) all place the emergency credential outside the SSO chain.
- Google removed standing operator access and measured a reliability dividend; GitLab keeps it for a small senior group and is still arguing with itself in public. The flip is your operational surface: many services favour Google's model, one deep application favours GitLab's.
- 'Believed unused' is not a security state: Cloudflare's four missed credentials, Microsoft's 2016 key and Okta's saved service account all show that a credential you cannot prove is dead is alive.
- Nobody in the corpus publishes their break-glass usage rate, the number that would reveal whether the third door is being overused; measuring it yourself is the only defence against the control degrading silently.
Scope
Why this, now. Just-in-time access and zero-standing-privilege moved from Google-scale practice into mid-size engineering orgs between 2022 and 2025 (Mercari, Figma, GitLab publishing openly), while the 2022-2024 breach reports (Uber, CircleCI, Okta, Cloudflare, Microsoft/CSRB) gave the first well-documented corpus of how human access actually fails in production.
What it does not cover. Product-side authorization for end users, service-to-service workload identity, the hour after a machine credential leaks, and the deploy pipeline as an attack surface, each covered by a separate dig in this collection. It does not cover any organisation's break-glass usage rate, which no source in the corpus publishes.
Other field guides
When the secret leaks: the hour after exposure
A field guide to the race that starts when a token goes public: who is allowed to kill a credential automatically, how fast the machinery actually wo…
26 sources · 18 organisations · 7 postmortemsThe leak goes around the tenant filter, not through it
A field guide to tenant isolation built from seven published cross-tenant incidents (Steam 2015, Cloudbleed 2017, GitHub 2021, ChaosDB 2021, AutoWarp…
27 sources · 25 organisations · 5 postmortemsThe breach is a missing check. The outage is the checker.
Reconstructs the authorization layer from the accounts of Google, Airbnb, Carta, Netflix, Figma, Slack and Gojek: relationship tuples, precomputed in…
26 sources · 24 organisations · 5 postmortems