Feature flags recurring outages
Four incidents in ten days; three shared one root cause, "nodes too small relative to pod resource requests", packing too many pods per node past 90 percent CPU. The fix raised requests and shrank the fleet.
Most containers use less than half of what they reserve, and the bill follows the reservation, not the usage. This guide reconstructs, from kernel commits, six years of Kubernetes pull requests, Google's own design records and two companies' 2025-26 incident reviews, why the gap exists, why the obvious way to close it (CPU limits) has its own failure record, and which machinery actually converts slack into money off the bill.
A reservation sized by fear, a bill sized by the reservation, and five kinds of machinery that attack the difference.
Strip the product names away and the problem reads like this: every tenant on shared compute declares in advance how much it needs. The platform packs machines by those declarations, buys hardware by the packing, and bills by the hardware. The declarations are written by people who are punished for guessing low (their service stalls at the worst moment) and never directly punished for guessing high (the overspend lands on a different team's dashboard, a quarter later). So the declarations are high, and they stay high, because nobody who could lower them is paid to.
The size of the result is measured, not folklore. Datadog's November 2023 container report, drawn from its customer fleet, found that over 65 percent of Kubernetes workloads use less than half of their requested CPU, and the same share holds for memory. CAST AI's 2024-2026 reports put average cluster CPU utilization near 10 percent; that host was unreachable from this session so the figure stands here as reported-but-unverified, and the verified Datadog number carries the argument on its own. The slack is not an accident of sloppy teams. It is the equilibrium of the incentive above.
Five kinds of machinery attack the gap, and each one has a public record of hurting someone. You can cap usage with limits, and the kernel's enforcement window becomes a tail latency problem with two kernel bugs in its history. You can shrink the declarations, and PostHog's October 2025 postmortem shows the failure mode of declarations that are too small. You can schedule other work into the gap, which is what Borg did for a 20 percent gain and what the Kubernetes QoS ladder was explicitly designed to import. You can repack and delete machines, which is the only step that converts any of this into money. Or you can pay a vendor to own the problem, which is what per-request billing actually is.
The reservation number is wrong in both directions at once. The measured fleet-wide story says requests are far too high (over 65 percent of containers use less than half, Datadog 2023). The freshest postmortem in this corpus says requests were too low: PostHog attributes three of its four October 2025 outages to "nodes too small relative to pod resource requests", and its remediation was to raise requests and run fewer pods per node. Both are true because a request does three jobs at once: it is the packing input, the performance insurance, and increasingly the billing line. The three jobs want three different numbers, and every team is quietly choosing which job wins.
Scope. This guide covers the economics and failure record of CPU and memory reservations on shared, orchestrated compute, with Kubernetes as the public record and Borg as its ancestry. It deliberately excludes GPU fleets (a different scarcity), the serverless tiers, VM-level rightsizing, reserved-instance and savings-plan financial engineering, and interruptible capacity, which has its own guide in this collection.
One declaration, two enforcement mechanisms, and three control loops. Every box below is traceable to a design record or shipped default.
The common shape across Borg, Kubernetes and the managed platforms built on them is a declaration with two numbers and an asymmetric contract. The request is what the scheduler packs by and what the system will defend under contention; the limit is a ceiling the kernel enforces. The Kubernetes QoS design document states the intent in one sentence: "when request < limit, the pod is guaranteed the request but can opportunistically scavenge the difference", and then states where the idea came from: "Borg increased utilization by about 20% when it started allowing use of such non-guaranteed resources, and we hope to see similar improvements in Kubernetes." The gap is not a defect of the model. The gap is the product; the design's whole purpose is to let someone else use it while defending your claim to take it back.
Enforcement is asymmetric because the two resources fail differently, and the project's own documentation is blunt about it: "cpu limits are enforced by CPU throttling ... a cpu limit is a hard limit the kernel enforces", while "memory limits are enforced by the kernel with out of memory (OOM) kills ... enforced reactively." CPU contention degrades you; memory contention kills someone, and the QoS ladder (Guaranteed, Burstable, BestEffort) exists to decide who. That asymmetry drives almost every decision in section 3: teams who would never drop a memory limit argue for years about dropping CPU limits, because the worst case of each is a different species.
Requests gate admission and set cpu.shares; limits set the CFS quota and the OOM score class. The ladder decides who is sacrificed when scavenged capacity is reclaimed: BestEffort first, Burstable by overage, Guaranteed last.
Recorded in: resource-qos.md, official docs source
A recommender consumes usage histograms and OOM events and proposes new requests. The community's version is named after Google's: the VPA design doc introduces itself as "aka rightsizing or autopilot". Until v1.35 every acted-on recommendation was an eviction, which is a large part of why adoption sat below 1 percent.
Recorded in: VPA proposal, KEP-1287, Datadog 2023
Shrinking requests saves nothing by itself; the fleet cost falls when a repacker drains a node and deletes it. Karpenter's design prices this exactly: "the cost savings is the price of the deleted node", weighed against a per-pod disruption cost, and replacement launches the cheaper node before draining the expensive one.
Recorded in: consolidation.md
Two divergence points separate the published systems. First, whether the ceiling is enforced at all: GitLab ships its own product's Helm chart with CPU requests set and the CPU limits block commented out, and Zalando's platform notes reason through the same trade; GKE's 2018 guidance assumes limits everywhere. Second, where the billing boundary sits: on node-billed platforms the cluster owner eats the slack, while GKE Autopilot bills "per second for vCPU, memory and disk resource requests", which moves the stranded capacity to Google and turns every inflated request into a visible line item. Google can afford to sell that contract because it has run the recommend loop internally for a decade; the buyer should notice that the vendor selling per-request billing is the one with the best rightsizing machinery on earth.
Three forks with recorded arguments, and the conditions that flip each one.
The limits argument is the best-documented fight in this corpus, and it was conducted in an issue tracker, not in blog posts. GitLab's 2020 internal thread shows how a serious platform team actually reasoned: not "limits are bad" but "there was a nasty bug in the Linux scheduler resulting in throttling/performance issues of all processes, even those not reaching their cgroup limits", followed by checking their GKE nodes' kernel against the fix, followed by the question that matters, "is it now (that the bug is fixed) safe to use CPU limits?" The honest answer from the kernel history is: safer, not safe. Chiluk's v5.4 commit removed the slice-expiration mechanism entirely because threaded applications could "hit a high percentage of periods throttled while simultaneously not consuming the allocated amount of quota". After v5.4 the pathology is gone, but the 100 ms accounting window remains, and a request-shaped burst inside one window is still a stall your average-CPU dashboard cannot see.
Decision 3 is worth reading as an institution rather than a feature. The ask was filed in November 2018, a one-line conditional by the reporter's own description. It took a kernel release, two closed-unmerged pull requests by different authors, one revival, and six years and three months to merge, and what finally landed is guarded by a feature gate so it can be turned off. Nothing about the idea was hard. What was hard was that a CPU limit means four different things (isolation, fairness, billing, prediction) to four different constituencies, and any change to its enforcement needs all four to agree. When you find a trivial-looking setting that has survived years of complaints, assume it is load-bearing for a constituency you have not met.
| Decision | Chosen | Rejected | Because | Evidence |
|---|---|---|---|---|
| CPU limits on latency path | Requests only (GitLab chart default) | Limits everywhere | 100 ms quota windows stall bursts; two kernel bugs throttled without overuse | chart defaults, #67577 |
| Memory ceiling | Always limited | Unlimited memory | Incompressible; the kernel kills rather than throttles, so the ladder must know who dies | QoS design |
| Utilization mechanism | Overcommit into the gap | Fleet-wide request shrinking | Borg's +20% is recorded; VPA adoption below 1% because resize meant eviction | QoS design, Datadog 2023 |
| Resize mechanics | In-place resize (stable v1.35) | Evict to resize | "For stateful or batch workloads, Pod restart is a serious disruption" | KEP-1287 |
| Quota on pinned pods | Exempt (v1.33, gated) | Quota always | Six-year argument; ceiling meant four things to four constituencies | #127525 |
| Who owns the slack | Cluster owner (node billing) | Vendor (per-request billing) | Autopilot charges "per second for vCPU, memory and disk resource requests"; good when requests are honest, expensive when inflated | GKE Autopilot, 2021 |
Four failure classes, each with a published incident and a rule you can carry into a design review.
The incidents in this corpus sort into four classes, and they bracket the topic neatly: two are failures of enforcing the ceiling, one is a failure of removing the slack, and one is a failure of believing the fleet average. None of them was detected by the metric the team was watching, which is the second pattern worth carrying: every one of these failure modes is invisible to an average-CPU dashboard.
Everything quantitative in the corpus, with its context and date. Measured, claimed and derived are marked.
| Metric | Value | At | Context | As of | Source |
|---|---|---|---|---|---|
| Containers using less than half their requested CPU (and memory) | >65% | Datadog customer fleet | Measured across billions of containers | Nov 2023 | Container Report |
| Utilization gained by scheduling into the gap | ~20% | Google Borg | Google's own figure, quoted in the Kubernetes QoS design to justify oversubscription | ~2015 | resource-qos.md |
| VPA adoption / HPA adoption | <1% / >50% | Datadog customer fleet | The rightsizer nobody runs vs the scaler half of everyone runs | Nov 2023 | Container Report |
| CFS enforcement window / slice / stranded per core | 100 ms / 5 ms / ~1 ms | Linux kernel | Per period; stranding bound stated in the fix commit | 2019 (v5.4) | de53fd7aed |
| Node CPU at which PostHog's cascade began | >90% | PostHog US fleet | Pool initialisation missed a 20 s startup timeout under this pressure | Oct 2025 | postmortem |
| Cumulative impact of the packing mistake | 14+ h | PostHog | Four incidents in ten days, error rates 34-97% | Oct 2025 | postmortem |
| PostHog's post-incident CPU alert threshold | 80% / 5 min | PostHog | Per-node and per-pod, paging | Oct 2025 | postmortem |
| GitLab webservice shipped defaults | 300m, no CPU limit | GitLab chart | Limits block present but commented out; sidekiq identical pattern | checked Oct 2026 | values.yaml |
| Quota exemption for pinned pods | v1.33 | Kubernetes | Gate DisableCPUQuotaWithExclusiveCPUs, default on; asked for Nov 2018 | Feb 2025 | #127525 |
| In-place pod resize | stable v1.35 | Kubernetes | Until then, acting on a rightsizing recommendation evicted the pod | 2026 | KEP-1287 |
| Google trace release built around this exact gap | 8 cells, 1 month, ~2.4 TiB | Google Borg | "The 2019 traces focus on resource requests and usage", with 5-minute usage histograms | May 2019 | ClusterData2019 |
The cost model, derived. Your effective price per used core is the list price per provisioned core divided by utilization. The arithmetic is unforgiving: at 50 percent utilization you pay double the list price per unit of work, at 10 percent you pay ten times it. Datadog's measured 65-percent-below-half figure means most teams sit past the doubling point on the workloads they thought they had sized. The model excludes the two things that justify some of the slack: the burst the service needs to boot and survive (PostHog paid 14 hours of outage for removing it), and the recovery headroom you draw on when a failure domain evacuates. Slack is only waste once those two are explicitly budgeted; the honest exercise is naming how much of your gap is insurance and how much is nobody's job to reclaim.
The Datadog figures are measured but come from one vendor's customer base, which skews toward teams that buy observability. The Borg 20 percent is Google's claim about its own system, made in a design document, not an audited measurement. CAST AI's widely quoted 8-13 percent average CPU utilization could not be fetched from this session's network and is deliberately absent from the table. The kernel numbers are defaults; both the period and the slice are tunable and some platforms tune them.
Every source behind this page, graded. The network allowlist for this research session reached GitHub, GitLab, Google Cloud and Datadog only; the gaps that creates are listed at the end of the section.
Four incidents in ten days; three shared one root cause, "nodes too small relative to pod resource requests", packing too many pods per node past 90 percent CPU. The fix raised requests and shrank the fleet.
A timeout cut from 1 s to 300 ms during database saturation; crash loops, retry amplification without circuit breakers, and probes that kept routing to dead pods for 45 minutes. 78 percent of US evaluations failed for 108 minutes.
An audit-event surge saturated one single-threaded Redis primary; deferral created a re-enqueue feedback loop; audit events dropped globally for just over three hours.
Opened by Ivan Babrou with a reproducer: quota throttles "well behaved tenants" that never consume their budget. Became the community's clearing-house for the limits argument for five years.
The drift check "can never hit", so quota expiration silently never ran. Fixing the bug made enforcement honest, and honest enforcement exposed the throttling everyone then reported.
Dave Chiluk's fix: threaded, non-CPU-bound apps were "throttled while simultaneously not consuming the allocated amount of quota" because up to 1 ms per core stranded every 100 ms period.
The original ask, framed by its reporter as a one-line conditional. Stayed open through two rejected pull requests and closed only when #127525 merged in 2025.
Died on a real objection: quota applied at the pod-level cgroup too, so the container-level exemption did not stop the throttling. The author closed it in favour of the kernel fix.
The second attempt, stalled on CRI side-effects across containerd and cri-o and the breadth of SIG review required, then closed for inactivity and revived by a third author.
The 2018 ask lands in v1.33 behind DisableCPUQuotaWithExclusiveCPUs, default on, with an explicit escape hatch for "users observing a regression".
GitLab's platform team re-derives the question from the kernel bug, checks its GKE kernels against the fix, and asks the honest question: is it now safe to use CPU limits?
The product GitLab ships to everyone else defaults to a 300m CPU request with the limits block present and commented out; the sidekiq chart repeats the pattern.
The published artifact of the EuroSys 2021 peak-prediction work: oversubscribe a machine by predicting its real peak from trace history instead of trusting summed requests. Built directly on the 2019 Borg traces.
The founding document: requests are guaranteed, the request-to-limit gap is scavengeable, nodes are deliberately oversubscribable, and Borg's 20 percent gain is the stated justification. CPU is throttled, never killed; the ladder decides memory.
Introduces itself as "aka rightsizing or autopilot": consume usage histograms and OOM events, recommend requests, apply them. The design's dependency on eviction is what the adoption number later indicted.
Names the blocker in its motivation: resources were immutable, so every resize was a recreation, "a serious disruption" for stateful and low-replica workloads. Now stable, which removes the main recorded excuse for not rightsizing.
The repack loop, priced: delete a node when its pods fit elsewhere, replace it with a cheaper one otherwise, score candidates by a per-pod disruption cost, and make "minimal changes to the cluster" per step.
A platform team's working notes, honest about their own uncertainty: "effect of CPU limits is not completely straight forward to understand", and higher limits than requests "allow to over-provision nodes, but has the danger of over-utilizing it".
Across Datadog's customer base: over 65 percent of Kubernetes workloads use less than half their requested CPU and memory; VPA adoption below 1 percent and flat since 2021; HPA above 50 percent.
The vendor framing of the contract: requests are guaranteed, "Kubernetes will make sure your containers get the CPU they requested and will throttle the rest", and the overcommitted state is presented as normal operation.
The billing boundary moves: "you pay only for the pods you use and you're billed per second for vCPU, memory and disk resource requests. No more worries about unused capacity!" The unused capacity still exists; Google now owns it.
Google's own framing of its public trace release: "the 2019 traces focus on resource requests and usage", with 5-minute CPU usage histograms rather than point samples, across eight cells for May 2019.
The enforcement asymmetry in the project's own words: CPU limits are "a hard limit the kernel enforces" by throttling; memory limits are enforced "reactively" by OOM kills that may lag the overage.
This session's network could not reach the canonical engineering-blog corpus (Uber's cpuset migration, Omio's throttling outage account, Indeed's two-part "Unthrottled", Buffer), the CAST AI utilization reports, any conference talk, or the paper PDFs for Borg (EuroSys 2015), Autopilot (EuroSys 2020) and the peak-prediction overcommit work (EuroSys 2021). Nothing from those sources is quoted here; where their claims matter, this guide cites the kernel commits, issue threads and artifact repositories that carry the same facts on reachable hosts. The widely circulated Autopilot slack figures are deliberately absent because they could not be verified against the paper. Those documents exist and are worth your time; the hunt queries in section 8 will find them from an unrestricted network.
Six rungs from reproducing the throttle to defending a utilization target with evidence.
On any Linux box with cgroups, run a multi-threaded, bursty workload under a CPU
quota (a Java or Go service doing periodic fan-out is ideal). Watch
cpu.stat: nr_periods, nr_throttled, throttled_time, against actual CPU
usage.
Done when: you can show throttled periods while total usage sits well under the quota, and explain it with the 100 ms window. Teaches: the ceiling is a latency instrument, not a capacity instrument.
For one real cluster, export requested vs used CPU and memory per workload over two weeks. Produce the two numbers that matter: total slack in cores, and the percentage of workloads under half their request (your private Datadog fact 5).
Done when: you can state "we reserve X cores and use Y at p95" with the query saved. Teaches: the gap is measurable in an afternoon; the hard part was never the measurement.
Pick a latency-sensitive, trusted service. Drop its CPU limit, keep its request and memory limit, and compare p99 and throttled-period counts for a week against a control. Check the node's other tenants for interference.
Done when: you have before/after p99 and nr_throttled, plus a neighbour-interference check, written down. Teaches: what the ceiling was costing you, and what (if anything) it was protecting.
Resize one deployment's requests from p95 steady-state usage plus its measured startup burst (time the pool initialisation and JIT warm-up under load). Then roll it and verify cold start still fits on a packed node.
Done when: a rolling restart under production load completes without startup-timeout crash loops. Teaches: PostHog's lesson at miniature scale; the first minute is the sizing constraint.
Enable a consolidation mechanism (Karpenter consolidation, or a manual drain-and- delete pass using the same arithmetic) and track node-hours per week before and after your rightsizing from rung 4.
Done when: the node count drops and the bill line follows; if requests shrank but nodes did not, you have proven the repack loop was the missing step. Teaches: where savings actually materialise.
On a staging cluster, pack a node past 85 percent CPU and roll a deployment onto it. Verify your alerts fire (PostHog settled on 80 percent sustained five minutes), your probes evict crash-looping pods quickly, and your retry budgets do not amplify.
Done when: the on-call is paged by CPU pressure before error rate, and the rollout degrades gracefully instead of cascading. Teaches: whether your margins are insurance or decoration.
The queries that found this material, grouped by what they surface. The vocabulary is the asset: throttling, CFS quota, bin-packing, slack, overcommit, rightsizing.
"CFS quota" throttling kernel bug kubernetes unthrottledrepo:kubernetes/kubernetes is:issue "cfs quotas" throttlingrepo:kubernetes/kubernetes is:pr is:closed is:unmerged cfs quotarepo:torvalds/linux commit "cfs_quota" throttling expirationpostmortem "cpu throttling" OR "noisy neighbor" production root causesite:github.com handbook post-mortems "resource requests"gitlab.com gl-infra "incident review" CPU saturation"cpu limits" site:gitlab.com gl-infra kuberneteskubernetes cluster utilization average CPU report waste requestscontainer report "requested" utilization datadog OR "cast ai"kubernetes "slack" requests usage "rightsizing" adoptionovercommitment datacenter "peak prediction" oversubscriptiongoogle cluster-data borg traces requests usage histogramsautopilot workload autoscaling google eurosys slack"avoiding cpu throttling" engineering cpuset containerized