Reliability & Consequence 30 September 2026 7 min read 1,636 words

Ask the node

Kubernetes has just made the containers in a running pod mutable, so that agents can spawn tool sandboxes in milliseconds. The record of which containers a node actually agreed to run was judged too expensive to store, and now lives on the node, ten statuses deep.

The argument

Kubernetes has moved the record of what a pod is actually running out of cluster state and onto the node, where nothing can watch it, list it, or keep it past ten containers.

The sentence worth sitting with is not in the proposal. It is in the justification for a small read-only endpoint bolted to the side of it.

On 25 September, after three and a half months of review, the Kubernetes project merged KEP-5972, Dynamic Containers. It does what a decade of Kubernetes documentation said could not be done: it lets main containers be added to and removed from a pod while the pod is running, through a new /dynamic subresource. .spec.containers is no longer immutable. The feature is targeted at alpha in v1.38, behind a DynamicContainers gate in both the API server and the kubelet, and three days later a follow-up pull request quietly added the production-readiness section that the first one had missed.

Because the desired container list and the running container list can now differ, the KEP adds a second, read-only subresource, /allocated, which returns what the node has actually agreed to run. Here is why it exists, in the authors' own words: the desired pod spec, stored in the pod resource, "can diverge from the allocated pod spec (stored locally by the Kubelet) for an arbitrary amount of time."

And here is where that second spec lives. It is "fetched directly from the Kubelet's /allocatedPods endpoint on-demand, thus avoiding additional storage overhead in the API server."

That is the whole argument of this piece. Kubernetes has made the running container set mutable, and in the same document decided that the authoritative record of what is running is too expensive to keep. So it has been moved onto the machine. Not as an oversight, not as a temporary alpha shortcut, but as a documented design choice with two rejected alternatives named underneath it.

Start with why anyone wants this, because the motivation is honest and it is not a small thing. The KEP's stated goals are a warm pool of pre-initialised pods, high-churn addition and removal of containers, and "sub-100ms latency required for interactive agentic workloads." The user stories name the buyers precisely: an agent or orchestrator that "spawns ephemeral tool-execution sandboxes on-the-fly within preallocated pods," and a two-tier scheduling arrangement in which Kubernetes is the macro-scheduler handing out a resource envelope while Ray or Slurm does the micro-scheduling inside it. The authors also considered the obvious alternative — just make pods start faster — and rejected it for a good reason: no amount of pod-startup optimisation removes the cost of populating a volume, pre-pulling a model or exchanging a credential before the workload can run. Pre-warm the envelope, then inject the work. As an engineering response to what agent runtimes actually need, this is correct.

The interesting part is what it costs, and the KEP is unusually candid about most of it. Third-party controllers that assume a static container array "may miss the workload entirely, panic due to index out-of-bounds errors, or fail to inject necessary sidecars." Short-lived containers at high frequency could "overwhelm the Kubelet's status manager, degrade kube-apiserver performance, and exhaust etcd write capacity." Both risks are named. Neither is the one that matters most.

The one that matters is that Kubernetes has never been primarily a scheduler. Plenty of things schedule work. What Kubernetes did that was durable was give a fleet a single, consistent, watchable account of itself — one place you could ask what exists, receive a stream of changes to it, and build compliance, cost attribution, image scanning and incident forensics as clients of that one record rather than as agents on every box. The API server's value was never that it held the desired state. It was that it held a state, completely, for everything, and would tell you when it changed.

/allocated is a read on one pod. It has no list. It has no watch. No informer can cache it, because the data is not in etcd to be watched — the request is proxied to the kubelet and answered by the node, on demand. Which means the three questions an operator actually asks now have three different answers. What was asked for? The API server, as before. What is running? Ask each node, one pod at a time, and only while the node is answering. What was running an hour ago? That one is worse.

When a container is removed and terminated, its status stays in the pod status — up to a point. The KEP sets the point: "Keep up to N (for alpha, N=10) removed container statuses: After more than N removed container statuses have accrued, delete the oldest status." Ten. In a feature whose first named goal is high-churn container addition and removal at sub-100ms, and whose flagship user story is an agent spawning tool sandboxes on the fly, the pod retains the last ten things it ran. The eleventh is gone from the API, and its logs go when container garbage collection takes them.

Ask the KEP how an operator is meant to know this feature is even in use and the answer completes the picture. Watch for container statuses in the Waiting state with reason Unallocated — or "monitor API Server audit logs for UPDATE events on the pods/dynamic subresource modifying .spec.containers." That is the sentence in which "what code ran on this machine" stops being cluster state and becomes a log line, in a facility that is not on by default, retained for as long as whoever configured it chose.

There is a precedent here, and the KEP cites it, which makes the departure deliberate rather than accidental. In-place pod resize faced the same divergence between desired and allocated, and solved it the other way: KEP-1287 put the allocated values into the API, at .status.containerStatuses[i].allocatedResources, and named a four-stage chain — desired, allocated, actuated, actual — so that each stage could be read from one place. KEP-5972 declines to extend that pattern, and says why. Mirroring the whole mutable spec into status "will drastically inflate the size of the pod object." Minting an AllocatedPod object per pod "would heavily inflate the API server object count and data size."

Both of those statements are true, and they are the strongest case against everything above. etcd is the known ceiling in large clusters; the KEP itself lists etcd write latency among the metrics that should trigger a rollback. Doubling the stored pod spec for every pod in every cluster, to serve a feature most clusters will never switch on, is a real cost borne by people who get no benefit from it. And this is not a careless proposal. Update permission on /dynamic is deliberately kept out of the default edit role and must be granted by a cluster admin. Privileged containers cannot be added. There is a kill switch, --disable-pod-subresources. Most tellingly, the authors invented a new fail-closed admission mechanism specifically to protect the installed base: a /dynamic request is rejected outright unless every admission webhook and policy that would have intercepted an equivalent create or update on pods is also registered on the subresource. Kubernetes has rarely gone to that much trouble to stop a new path from quietly bypassing an old guard.

Which is exactly what makes the asymmetry legible. Confronted with the risk that policy would be evaded, the project built a new enforcement primitive. Confronted with the risk that the record would thin out, it reached for garbage collection and set N to ten. The gate got a mechanism; the account got a retention limit. Even the new mechanism is shaped by the same instinct: it honours a handler's MatchConstraints and selectors when deciding whether coverage exists, but MatchConditions "are not evaluated as part of this decision" — it checks that something is registered, not that it would have looked.

The readiness rule shows where this ends up. A container added but not yet allocated is Waiting with reason Unallocated, and the KEP updates pod readiness so that such containers "are ignored when computing pod readiness." The reasoning is sound; a deferred allocation should not pull a serving pod out of its Service. The consequence is that a pod can be Ready, and stay Ready for an arbitrary amount of time, while its own spec declares a container that no node has agreed to run. The narrower fix — letting containers be exempted from readiness properly — is deferred to a Kubernetes issue that has been open since September 2024.

None of this is a plea for immutability. The envelope model is the right shape for agent workloads, and the KEP earns its scope. But notice what is being traded and for whom. The latency does not exist yet: against the stated sub-100ms goal, the alpha SLO is a median under 500ms and a 95th percentile under five seconds, "heavily constrained by containerd/runc latency." The tracking issue still carries the ambition in its title — Optimistic Execution & Hierarchical Resource Delegation — while the merged KEP lists bypassing the API server as an explicit non-goal. So the record is being economised now, in advance, for a speed the implementation has not reached, on behalf of software that writes its own next container.

That is the part to hold onto. The more autonomous the thing inside the pod, the more the account of what it ran is the only thing anyone can be held to. We are building envelopes for processes that decide for themselves what to execute, and simultaneously deciding that keeping track of what they executed does not scale. A pod that spawns a thousand sandboxes in a day will remember ten. Ask it what it ran, and it will tell you about this afternoon.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. KEP-5972 Dynamic Containers README Kubernetes Enhancements · 2026-09-25
  2. KEP-5972 kep.yaml Kubernetes Enhancements · 2026-06-07
  3. KEP-5972: Dynamic Containers (PR #6169) Kubernetes Enhancements · 2026-09-25
  4. KEP-5972: Dynamic Containers - add missed PRR section (PR #6435) Kubernetes Enhancements · 2026-09-28
  5. Dynamic Pod Mutation: Optimistic Execution & Hierarchical Resource Delegation (issue #5972) Kubernetes Enhancements · 2026-09-25
  6. KEP-1287 In-place Update of Pod Resources README Kubernetes Enhancements · 2026-09-25
  7. Idea: Pod-level probes or exclude some containers from pod readiness (issue #127276) Kubernetes · 2024-09-10

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

kubernetespod lifecycleauditagent sandboxescontrol plane