advanced 4 min answer

A new telemetry service is rolled out across a large Kubernetes fleet to improve visibility. Within minutes, DNS-based service discovery degrades across the cluster and a substantial outage follows, lasting hours. OpenAI published this for 11 December 2024. Reason backwards: what failed, which design decision made it possible, and why was recovery slow rather than a simple rollback?

openaikubernetescontrol planetelemetryblast radiusrecovery
Show the full answer Hide the answer

The trigger

A new telemetry service was deployed cluster-wide. OpenAI's published account is that its configuration caused resource-intensive Kubernetes API operations, and the cost of those operations scaled with the size of the cluster — so the change was tested successfully on smaller clusters and only became fatal at full scale. The API servers saturated.

The immediate lesson is narrow and important: an observability component is a client of the control plane, and a client that lists or watches cluster-wide resources imposes a load proportional to the fleet. Deployed to every node at once, it multiplies that by the node count. The change was additive and read-only, which is exactly the profile of change that review treats as safe.

Why it propagated

The control plane is not a management convenience; it is a runtime dependency for the things built on top of it. In this case, DNS-based service discovery resolved through machinery that depends on the Kubernetes API. When the API servers degraded, service discovery degraded, and services that were running perfectly well became unreachable because nothing could resolve where they were.

This is the coupling that matters: the data plane was healthy. No application crashed. The failure was entirely in the layer that tells healthy things how to find each other, and from the outside that is indistinguishable from everything being down.

Why recovery lagged

Here is the part that turns a ten-minute mistake into a multi-hour outage, and it is the generalisable lesson.

The remediation — removing the telemetry service — required the Kubernetes control plane, which was the saturated resource. To fix it, engineers needed to issue API calls to the API servers that were too overloaded to serve API calls. The system had entered a state where the mechanism for changing it was downstream of the thing that was broken.

This is a deadlock, not a slowdown. It does not resolve by waiting, because the load is being generated by workloads that are still running and still have no reason to stop. Breaking it needs out-of-band action — scaling the control plane, blocking the offending clients at the admission or network layer, or reaching individual nodes directly — all of which are slower and riskier than the normal path, and none of which anyone practises.

The structural fix versus the tempting local fix

The tempting local fix: test the telemetry service on a large cluster before rolling it out. Correct, insufficient, and it generalises to nothing. The next incident will be a different component with a different scaling behaviour.

The structural fixes, in order of value:

  1. A break-glass path that does not traverse the broken layer. The single highest-value item. There must be a way to remove a workload that does not require the control plane to be healthy — direct node access, a static manifest path, an independent emergency control plane. The time to build this is not during the incident, and the question to ask of any shared dependency is "if this is saturated, how do I change it?"
  2. Resource isolation on the control plane. Rate limits and priority classes so that no client can starve the API servers, and so that administrative traffic outranks workload traffic. Kubernetes has API Priority and Fairness for exactly this; it is off the critical path of most teams' attention until it is the only thing that matters.
  3. Staged rollout with a blast-radius-aware stop. Not percentage-of-fleet but percentage-of-fleet-per-cluster, with a bake time long enough for load effects to appear. The flaw in a fast cluster-wide rollout is that the effect being guarded against is superlinear in deployment breadth, so a 5% canary shows nothing at all.
  4. Decouple service discovery from the control plane's health, with cached or static fallback resolution, so that a control-plane degradation stops being an immediate data-plane outage.

When not to generalise this

The break-glass work above is proportionate to a fleet whose control plane is a hard runtime dependency for service discovery. On a smaller cluster, or one where services address each other through a load balancer or static configuration rather than through control-plane-backed discovery, a saturated API server is a deployment freeze rather than an outage — unpleasant, not customer-visible, and not worth an independent emergency control plane. Build the escape hatch when the control plane is on the request path, and when it is not, spend the effort on the rate limits instead, which are cheap and help regardless.

The general lesson

Monitoring is not exempt. Observability changes arrive with an implicit presumption of safety — they only read, they only report, they exist to make things better — and they are deployed everywhere at once by their nature. That combination makes them one of the most effective ways to cause a correlated failure.

The question to put to any fleet-wide agent before it ships: what does one instance cost the shared control plane, and what is that multiplied by every node? If nobody can answer, the rollout is not ready, however read-only the change is.