On 11 December 2024 OpenAI deployed a new telemetry service into its Kubernetes clusters. Its configuration caused every node to perform Kubernetes API operations whose cost scaled with the size of the cluster; API servers saturated and DNS-based service discovery failed across most large clusters, with services unavailable from 15:16 to 19:38 PST. The change had been tested in staging. What structurally failed, what could staging not test, and where would copying their response be a mistake?
Show the full answer Hide the answer
The situation they were in
A platform agent on every node is the most privileged tenant of a shared control plane, and it is usually not treated as a tenant at all. Product workloads get namespaces, quotas and limits; the platform's own DaemonSets get cluster-wide read access and no budget.
What structurally failed
The load was quadratic in cluster size, and nobody had a number for it. One node performing an operation whose cost scales with the number of objects in the cluster is cheap. N nodes each performing it is N × O(N) work on the API servers. At 50 nodes that product is 2,500 units; at 4,000 nodes it is 16 million. The same configuration is harmless at one scale and fatal at the other, with no warning in between and no alert that fires on "cost per call × callers".
The second structural fact is what the data plane depends on. Running pods keep serving without the control plane, which is why people call the data plane independent. In-cluster DNS is not independent: resolution is backed by control-plane state, so when the API servers went down, name resolution went with them, and a management-component problem became a total traffic failure.
Why staging could not see it
Staging reproduced the configuration. It could not reproduce the quadratic term, because the term is the cluster. A staging cluster two orders of magnitude smaller tests correctness and tests nothing about control-plane cost.
Worse, the rollout's own health signal was wrong. DNS answers were already cached, so name resolution kept working for a while after the control plane degraded, the canary looked healthy and the rollout continued. A gate that reads a cached dependency is a gate that approves the change that is breaking it.
What it cost them
Roughly three hours of degradation across the public services, and a slow recovery, because remediation needed the resource that was saturated. Fixing a misbehaving DaemonSet means talking to the API server, which was the thing under load. Any plan whose first step is "roll back the deployment" assumes the control plane answers.
What to change, in order
- Give platform agents an API budget: Kubernetes API Priority and Fairness classifies and caps a DaemonSet's share so one agent cannot starve everything else, and admin requests keep a reserved seat.
- Make agents watch, not list. An informer with a resourceVersion watch costs one initial list plus deltas. Repeated full lists from every node is what produces the N × O(N) shape.
- Load-test against cluster shape, not cluster correctness. A synthetic cluster with production object counts, built with fake nodes, is the only place the quadratic term appears.
- Pick a health signal with no cache in front of it for control-plane rollouts: API server request latency and inflight counts, not DNS success.
Where copying this would be a mistake
On a 40-node cluster, an agent that lists everything every 30 seconds costs nothing measurable, and fairness tuning, a shape-accurate load environment and an admin-reserved access path are perhaps a quarter of an engineer's year for a risk that does not exist. The thresholds that make this real: more than a few hundred nodes, more than one cluster-wide agent, or a control plane you cannot replace in minutes. Below that, the lesson is one line in a review checklist - does this agent's work grow with the cluster.
Common weak answers
- "They should have tested in staging." They did. The missing test was cluster shape, and no amount of functional staging produces it.
- "Canary the rollout." There was a staged rollout. It was approved by a signal with a cache in front of it, which is the actual lesson.
- "A cluster per team." That moves the cost and multiplies the control planes you patch. Quotas are cheaper until tenant criticality genuinely differs.