Requested vs used  / field guide
Practitioner field guide · 2026-10-03 · cost & efficiency

The compute you reserve but never use

Most containers use less than half of what they reserve, and the bill follows the reservation, not the usage. This guide reconstructs, from kernel commits, six years of Kubernetes pull requests, Google's own design records and two companies' 2025-26 incident reviews, why the gap exists, why the obvious way to close it (CPU limits) has its own failure record, and which machinery actually converts slack into money off the bill.

23 primary sources 12 organisations 6 published incidents Evidence through May 2026 Read: 29 min
01

The territory

A reservation sized by fear, a bill sized by the reservation, and five kinds of machinery that attack the difference.

Strip the product names away and the problem reads like this: every tenant on shared compute declares in advance how much it needs. The platform packs machines by those declarations, buys hardware by the packing, and bills by the hardware. The declarations are written by people who are punished for guessing low (their service stalls at the worst moment) and never directly punished for guessing high (the overspend lands on a different team's dashboard, a quarter later). So the declarations are high, and they stay high, because nobody who could lower them is paid to.

The size of the result is measured, not folklore. Datadog's November 2023 container report, drawn from its customer fleet, found that over 65 percent of Kubernetes workloads use less than half of their requested CPU, and the same share holds for memory. CAST AI's 2024-2026 reports put average cluster CPU utilization near 10 percent; that host was unreachable from this session so the figure stands here as reported-but-unverified, and the verified Datadog number carries the argument on its own. The slack is not an accident of sloppy teams. It is the equilibrium of the incentive above.

>65%
of containers use less than half of their requested CPU and memory
+20%
utilization Borg gained by scheduling other work into the gap
<1%
of organisations run the Vertical Pod Autoscaler, flat since 2021
6.3 yr
from the first ask to stop throttling pinned pods (Nov 2018) to the merged fix (Feb 2025)

Five kinds of machinery attack the gap, and each one has a public record of hurting someone. You can cap usage with limits, and the kernel's enforcement window becomes a tail latency problem with two kernel bugs in its history. You can shrink the declarations, and PostHog's October 2025 postmortem shows the failure mode of declarations that are too small. You can schedule other work into the gap, which is what Borg did for a 20 percent gain and what the Kubernetes QoS ladder was explicitly designed to import. You can repack and delete machines, which is the only step that converts any of this into money. Or you can pay a vendor to own the problem, which is what per-request billing actually is.

The finding that reframes the topic

The reservation number is wrong in both directions at once. The measured fleet-wide story says requests are far too high (over 65 percent of containers use less than half, Datadog 2023). The freshest postmortem in this corpus says requests were too low: PostHog attributes three of its four October 2025 outages to "nodes too small relative to pod resource requests", and its remediation was to raise requests and run fewer pods per node. Both are true because a request does three jobs at once: it is the packing input, the performance insurance, and increasingly the billing line. The three jobs want three different numbers, and every team is quietly choosing which job wins.

Figure 1 · The bill follows the declaration, not the usage

declares

sum fills nodes

over 65% use less than
half of requested

only a repacker turns
this into savings

Service owner
sizes for the worst hour

Pod requests

Scheduler bin-packing

Node count

The bill

Actual usage
(the average minute)

Slack

declares

sum fills nodes

over 65% use less than
half of requested

only a repacker turns
this into savings

Service owner
sizes for the worst hour

Pod requests

Scheduler bin-packing

Node count

The bill

Actual usage
(the average minute)

Slack

Node count, and therefore spend, is a function of summed requests; actual usage only touches the bill if a repacking loop shrinks the node count. Sources: Kubernetes QoS design, Datadog 2023, Karpenter consolidation design.
Diagram source

Scope. This guide covers the economics and failure record of CPU and memory reservations on shared, orchestrated compute, with Kubernetes as the public record and Borg as its ancestry. It deliberately excludes GPU fleets (a different scarcity), the serverless tiers, VM-level rightsizing, reserved-instance and savings-plan financial engineering, and interruptible capacity, which has its own guide in this collection.

02

How the gap is actually managed

One declaration, two enforcement mechanisms, and three control loops. Every box below is traceable to a design record or shipped default.

The common shape across Borg, Kubernetes and the managed platforms built on them is a declaration with two numbers and an asymmetric contract. The request is what the scheduler packs by and what the system will defend under contention; the limit is a ceiling the kernel enforces. The Kubernetes QoS design document states the intent in one sentence: "when request < limit, the pod is guaranteed the request but can opportunistically scavenge the difference", and then states where the idea came from: "Borg increased utilization by about 20% when it started allowing use of such non-guaranteed resources, and we hope to see similar improvements in Kubernetes." The gap is not a defect of the model. The gap is the product; the design's whole purpose is to let someone else use it while defending your claim to take it back.

Enforcement is asymmetric because the two resources fail differently, and the project's own documentation is blunt about it: "cpu limits are enforced by CPU throttling ... a cpu limit is a hard limit the kernel enforces", while "memory limits are enforced by the kernel with out of memory (OOM) kills ... enforced reactively." CPU contention degrades you; memory contention kills someone, and the QoS ladder (Guaranteed, Burstable, BestEffort) exists to decide who. That asymmetry drives almost every decision in section 3: teams who would never drop a memory limit argue for years about dropping CPU limits, because the worst case of each is a different species.

Figure 2 · Reference architecture: the three loops around one declaration

resize; an eviction until
in-place landed (stable v1.35)

delete or replace nodes;
'cost savings is the price
of the deleted node'

Declaration:
requests + limits

Scheduler packs
nodes by requests

Kernel enforcement:
shares for requests,
CFS quota for limits

Telemetry: usage histograms,
nr_throttled, OOM events

Recommender
(VPA, 'aka autopilot')

Repacker / consolidation

resize; an eviction until
in-place landed (stable v1.35)

delete or replace nodes;
'cost savings is the price
of the deleted node'

Declaration:
requests + limits

Scheduler packs
nodes by requests

Kernel enforcement:
shares for requests,
CFS quota for limits

Telemetry: usage histograms,
nr_throttled, OOM events

Recommender
(VPA, 'aka autopilot')

Repacker / consolidation

The enforce loop runs every 100 ms in the kernel; the recommend loop runs over days of telemetry; the repack loop is the only one that touches the bill. Reconstructed from the QoS design, the VPA proposal, Karpenter's consolidation design and Google's trace documentation.
Diagram source

The declaration and the QoS ladder

Requests gate admission and set cpu.shares; limits set the CFS quota and the OOM score class. The ladder decides who is sacrificed when scavenged capacity is reclaimed: BestEffort first, Burstable by overage, Guaranteed last.

Recorded in: resource-qos.md, official docs source

The recommend loop

A recommender consumes usage histograms and OOM events and proposes new requests. The community's version is named after Google's: the VPA design doc introduces itself as "aka rightsizing or autopilot". Until v1.35 every acted-on recommendation was an eviction, which is a large part of why adoption sat below 1 percent.

Recorded in: VPA proposal, KEP-1287, Datadog 2023

The repack loop

Shrinking requests saves nothing by itself; the fleet cost falls when a repacker drains a node and deletes it. Karpenter's design prices this exactly: "the cost savings is the price of the deleted node", weighed against a per-pod disruption cost, and replacement launches the cheaper node before draining the expensive one.

Recorded in: consolidation.md

Two divergence points separate the published systems. First, whether the ceiling is enforced at all: GitLab ships its own product's Helm chart with CPU requests set and the CPU limits block commented out, and Zalando's platform notes reason through the same trade; GKE's 2018 guidance assumes limits everywhere. Second, where the billing boundary sits: on node-billed platforms the cluster owner eats the slack, while GKE Autopilot bills "per second for vCPU, memory and disk resource requests", which moves the stranded capacity to Google and turns every inflated request into a visible line item. Google can afford to sell that contract because it has run the recommend loop internally for a decade; the buyer should notice that the vendor selling per-request billing is the one with the best rightsizing machinery on earth.

03

The decisions that matter

Three forks with recorded arguments, and the conditions that flip each one.

Decision 1: do latency-sensitive services get CPU limits at all?

Chosen (by the operators on record)
  • Requests only, no CPU limit: GitLab's shipped chart defaults (300m request, limits commented out); the position argued by Cloudflare's Ivan Babrou in kubernetes #67577 since 2018
  • Memory limits stay on; nobody on record argues against those
Rejected
  • Limits everywhere for predictability, the 2018 GKE guidance: the quota is enforced in 100 ms windows, so a bursty service stalls inside a window even when its average is far below the limit
  • Two kernel bugs (fixed v4.18 and v5.4) meant years of throttling that tracked no real overuse at all
Flips when
  • Neighbours are untrusted and interference costs more than throttling
  • The platform bills or charges back by the ceiling
  • You can pin instead: whole-core, latency-critical shapes go to Decision 3

The limits argument is the best-documented fight in this corpus, and it was conducted in an issue tracker, not in blog posts. GitLab's 2020 internal thread shows how a serious platform team actually reasoned: not "limits are bad" but "there was a nasty bug in the Linux scheduler resulting in throttling/performance issues of all processes, even those not reaching their cgroup limits", followed by checking their GKE nodes' kernel against the fix, followed by the question that matters, "is it now (that the bug is fixed) safe to use CPU limits?" The honest answer from the kernel history is: safer, not safe. Chiluk's v5.4 commit removed the slice-expiration mechanism entirely because threaded applications could "hit a high percentage of periods throttled while simultaneously not consuming the allocated amount of quota". After v5.4 the pathology is gone, but the 100 ms accounting window remains, and a request-shaped burst inside one window is still a stall your average-CPU dashboard cannot see.

Decision 2: close the gap by shrinking requests, or schedule into it?

Chosen (by the systems with the gains on record)
  • Schedule into the gap: Borg's non-guaranteed resources bought "about 20%" utilization, and the Kubernetes QoS ladder was designed to import exactly that
  • Google's overcommit line of work predicts the machine's real peak instead of trusting summed requests; its published simulator is built on the 2019 traces
Rejected (by revealed preference)
  • Fleet-wide request shrinking via VPA: below 1 percent adoption, flat since 2021, while HPA passes 50 percent
  • The recorded reasons: acting on a recommendation was an eviction until in-place resize went stable (v1.35), and VPA fights HPA on the same metric
Flips when
  • In-place resize removes the eviction cost (it is now stable; the excuse is expiring)
  • Billing is per-request (Autopilot): shrinking requests becomes the whole game
  • The workload is a singleton or stateful: overcommit's reclaim lands on it hardest, so size it honestly instead

Decision 3: do pinned, whole-core pods still get a quota?

Chosen, eventually
  • Exempt them: pods qualifying for static CPU placement "will not have cfs quota enforced", merged February 2025, Kubernetes v1.33, feature gate DisableCPUQuotaWithExclusiveCPUs, default on
Rejected twice first
  • 2019, PR #75682: quota still applied at the pod-level cgroup; author closed it, "fully supportive of closing this PR in favour of the kernel fix"
  • 2022-23, PR #107589: stalled on CRI side-effects and SIG review breadth, closed for inactivity, revived by a third author
Flips when
  • The ceiling carries meaning beyond isolation (billing, chargeback, multi-tenant fairness): then even a pinned pod keeps its quota
  • You are below v1.33: pinning does not stop throttling, plan around it

Decision 3 is worth reading as an institution rather than a feature. The ask was filed in November 2018, a one-line conditional by the reporter's own description. It took a kernel release, two closed-unmerged pull requests by different authors, one revival, and six years and three months to merge, and what finally landed is guarded by a feature gate so it can be turned off. Nothing about the idea was hard. What was hard was that a CPU limit means four different things (isolation, fairness, billing, prediction) to four different constituencies, and any change to its enforcement needs all four to agree. When you find a trivial-looking setting that has survived years of complaints, assume it is load-bearing for a constituency you have not met.

Figure 3 · One setting, six years: the quota-exemption arc

Nov 2018: #70585 filed,
'just set cpuQuota=-1'

2019: PR #75682 closed unmerged,
pod-level cgroup still throttles

Nov 2019: kernel v5.4
fixes slice stranding

2022-23: PR #107589 closed unmerged,
CRI side-effects, review breadth

Feb 2025: #127525 merged,
v1.33, gated, default on

Nov 2018: #70585 filed,
'just set cpuQuota=-1'

2019: PR #75682 closed unmerged,
pod-level cgroup still throttles

Nov 2019: kernel v5.4
fixes slice stranding

2022-23: PR #107589 closed unmerged,
CRI side-effects, review breadth

Feb 2025: #127525 merged,
v1.33, gated, default on

Each failed attempt recorded a different constituency's objection; the merged fix still ships behind an escape-hatch gate. Sources: #70585, #75682, #107589, #127525.
Diagram source

Figure 4 · The decision path, as the record supports it

yes

no

yes

no

no

yes

Untrusted neighbours, or the
ceiling carries billing meaning?

Keep CPU limits.
Alert on nr_throttled,
not on average CPU

Whole-core, latency-critical
shape?

Guaranteed QoS, integer CPUs,
static policy; quota off
from v1.33

Requests only.
Memory limits stay on.
Alert on node CPU pressure

Is freed request-space
becoming fewer nodes?

Add the repack loop:
consolidation prices savings
against disruption

Hold. Re-size requests against
the boot peak, not the mean

yes

no

yes

no

no

yes

Untrusted neighbours, or the
ceiling carries billing meaning?

Keep CPU limits.
Alert on nr_throttled,
not on average CPU

Whole-core, latency-critical
shape?

Guaranteed QoS, integer CPUs,
static policy; quota off
from v1.33

Requests only.
Memory limits stay on.
Alert on node CPU pressure

Is freed request-space
becoming fewer nodes?

Add the repack loop:
consolidation prices savings
against disruption

Hold. Re-size requests against
the boot peak, not the mean

Terminal nodes are actions with sources behind them, not preferences. The left branch is Decision 1, the middle is Decision 3, the bottom is the repack loop from Karpenter's design.
Diagram source
DecisionChosenRejectedBecauseEvidence
CPU limits on latency pathRequests only (GitLab chart default)Limits everywhere100 ms quota windows stall bursts; two kernel bugs throttled without overusechart defaults, #67577
Memory ceilingAlways limitedUnlimited memoryIncompressible; the kernel kills rather than throttles, so the ladder must know who diesQoS design
Utilization mechanismOvercommit into the gapFleet-wide request shrinkingBorg's +20% is recorded; VPA adoption below 1% because resize meant evictionQoS design, Datadog 2023
Resize mechanicsIn-place resize (stable v1.35)Evict to resize"For stateful or batch workloads, Pod restart is a serious disruption"KEP-1287
Quota on pinned podsExempt (v1.33, gated)Quota alwaysSix-year argument; ceiling meant four things to four constituencies#127525
Who owns the slackCluster owner (node billing)Vendor (per-request billing)Autopilot charges "per second for vCPU, memory and disk resource requests"; good when requests are honest, expensive when inflatedGKE Autopilot, 2021
04

What broke in production

Four failure classes, each with a published incident and a rule you can carry into a design review.

The incidents in this corpus sort into four classes, and they bracket the topic neatly: two are failures of enforcing the ceiling, one is a failure of removing the slack, and one is a failure of believing the fleet average. None of them was detected by the metric the team was watching, which is the second pattern worth carrying: every one of these failure modes is invisible to an average-CPU dashboard.

Kernel record

Class 1: the ceiling bites before the budget is spent

AssumptionThrottling means the container used its quota.
What happenedBefore v4.18, a clock-drift bug in expire_cfs_rq_runtime "can never hit", so stale quota lingered; the fix exposed the second bug, in which per-CPU slices strand "typically no more than min_cfs_rq_runtime=1ms per cpu" per 100 ms period. On an 88-core box that is most of a small container's budget. Threaded apps hit "a high percentage of periods throttled while simultaneously not consuming the allocated amount of quota".
Blast radiusEvery container platform enforcing CPU limits on kernels v4.18 through v5.3, roughly 2018-2020; surfaced publicly through kubernetes #67577, opened by Cloudflare's Ivan Babrou in August 2018.
FixKernel commit de53fd7aed (v5.4) removed slice expiration; operators meanwhile removed CPU limits or disabled CFS quota in the kubelet.
Design rulenr_throttled over nr_periods is the signal, never average CPU. And a conclusion drawn from five years of buggy enforcement ("limits are poison") needs re-testing on a fixed kernel before it becomes your policy.
Postmortem

Class 2: the slack was armor, and it was removed

AssumptionRequests sized for steady state are safe to pack tightly; density is pure efficiency.
What happenedPostHog's nodes were "too small relative to pod resource requests, causing Kubernetes to pack too many pods per node and saturate CPU capacity" past 90 percent. Under that pressure, new pods could not finish Postgres pool initialisation (TLS, auth) inside the 20-second startup timeout, crash-looped, and Envoy retries fanned each degraded request into dozens of concurrent Redis reads, overwhelming a cache shared with the main application.
Blast radiusThree of four incidents in ten days of October 2025; over 14 hours of cumulative major impact; error rates between 34 and 97 percent; one episode ran 7 hours 9 minutes because "CPU alerting was completely absent".
FixRaised pod requests and re-sized the fleet: "running smaller fleets with better-resourced pods rather than larger fleets with CPU-bound pods", plus CPU alerts at 80 percent sustained for five minutes, plus Envoy retry limits.
Design ruleSize requests for the startup peak, not the steady state; a pod's most expensive minute is its first. And expect CPU saturation to present as dependency failures: PostHog's dashboards blamed Redis and Postgres while both "confirmed they were healthy".
Postmortem

Class 3: slow becomes down at the timeout boundary

AssumptionA tighter timeout is a tuning detail, separable from capacity margins.
What happenedA deployment cut PostHog's database connection timeout from 1 s to 300 ms while the writer database was already saturated by ingestion. Pool creation could no longer complete, pods crash-looped, retry logic without circuit breakers amplified the load, and misconfigured probes meant "Kubernetes continued routing traffic to pods in crash loops for up to 45 minutes".
Blast radius108 minutes on September 29, 2025; roughly 78 percent of US flag evaluations returned 504s, including read-only paths that never needed the writer.
FixTimeouts moved to runtime configuration and raised to cover peak-load pool initialisation; circuit breakers; probes that actually evict crash-looping pods.
Design ruleTimeouts, retry budgets and capacity margins are one budget, not three. Any timeout must cover the operation's cost under CPU and dependency pressure, which is when it will actually be exercised.
Postmortem

Class 4: the average hides the saturated core

AssumptionFleet-level headroom protects every component; utilization is a fleet-level number.
What happenedA surge of audit-event jobs saturated the CPU of one Redis primary at GitLab. "Redis's single-threaded architecture limited CPU scaling, compounding the resource bottleneck", and the mitigation (deferring jobs) "appears to have created a feedback loop" in which Sidekiq's own re-enqueues became the dominant load.
Blast radiusAudit events dropped globally for 3 hours 7 minutes on May 19, 2026; a near-miss repeat of an S1 the week before that had broken cluster quorum.
FixDropped the jobs outright instead of deferring, then a feature change lock on the offending worker while corrective work landed.
Design ruleUtilization targets are per-bottleneck, not per-fleet. A cluster at 10 percent still contains single-threaded components whose one core is the only capacity that matters; the fleet average is not a planning number for them.

Figure 5 · How tight packing became a 123-minute outage (PostHog, October 28, 2025)

Shared RedisPostgresNew podPacked node (CPU >90%)SchedulerRoutine deployShared RedisPostgresNew podPacked node (CPU >90%)SchedulerRoutine deploymisses 20 s startuptimeoutwrite storm, keyevictionsrollout triggers flags podsplace by requests (node alreadytight)start under CPU pressureinitialise pool (TLS, auth)crash loop, capacity fallscache misses, synchronous full-state writesmain application degrades too
Shared RedisPostgresNew podPacked node (CPU >90%)SchedulerRoutine deployShared RedisPostgresNew podPacked node (CPU >90%)SchedulerRoutine deploymisses 20 s startuptimeoutwrite storm, keyevictionsrollout triggers flags podsplace by requests (node alreadytight)start under CPU pressureinitialise pool (TLS, auth)crash loop, capacity fallscache misses, synchronous full-state writesmain application degrades too
The triggering deployment contained no flags-service changes at all; the packed node turned a routine rollout into the incident. Reconstructed from PostHog's postmortem.
Diagram source
05

Numbers you can plan against

Everything quantitative in the corpus, with its context and date. Measured, claimed and derived are marked.

MetricValueAtContextAs ofSource
Containers using less than half their requested CPU (and memory)>65%Datadog customer fleetMeasured across billions of containersNov 2023Container Report
Utilization gained by scheduling into the gap~20%Google BorgGoogle's own figure, quoted in the Kubernetes QoS design to justify oversubscription~2015resource-qos.md
VPA adoption / HPA adoption<1% / >50%Datadog customer fleetThe rightsizer nobody runs vs the scaler half of everyone runsNov 2023Container Report
CFS enforcement window / slice / stranded per core100 ms / 5 ms / ~1 msLinux kernelPer period; stranding bound stated in the fix commit2019 (v5.4)de53fd7aed
Node CPU at which PostHog's cascade began>90%PostHog US fleetPool initialisation missed a 20 s startup timeout under this pressureOct 2025postmortem
Cumulative impact of the packing mistake14+ hPostHogFour incidents in ten days, error rates 34-97%Oct 2025postmortem
PostHog's post-incident CPU alert threshold80% / 5 minPostHogPer-node and per-pod, pagingOct 2025postmortem
GitLab webservice shipped defaults300m, no CPU limitGitLab chartLimits block present but commented out; sidekiq identical patternchecked Oct 2026values.yaml
Quota exemption for pinned podsv1.33KubernetesGate DisableCPUQuotaWithExclusiveCPUs, default on; asked for Nov 2018Feb 2025#127525
In-place pod resizestable v1.35KubernetesUntil then, acting on a rightsizing recommendation evicted the pod2026KEP-1287
Google trace release built around this exact gap8 cells, 1 month, ~2.4 TiBGoogle Borg"The 2019 traces focus on resource requests and usage", with 5-minute usage histogramsMay 2019ClusterData2019

The cost model, derived. Your effective price per used core is the list price per provisioned core divided by utilization. The arithmetic is unforgiving: at 50 percent utilization you pay double the list price per unit of work, at 10 percent you pay ten times it. Datadog's measured 65-percent-below-half figure means most teams sit past the doubling point on the workloads they thought they had sized. The model excludes the two things that justify some of the slack: the burst the service needs to boot and survive (PostHog paid 14 hours of outage for removing it), and the recovery headroom you draw on when a failure domain evacuates. Slack is only waste once those two are explicitly budgeted; the honest exercise is naming how much of your gap is insurance and how much is nobody's job to reclaim.

Read these carefully

The Datadog figures are measured but come from one vendor's customer base, which skews toward teams that buy observability. The Borg 20 percent is Google's claim about its own system, made in a design document, not an audited measurement. CAST AI's widely quoted 8-13 percent average CPU utilization could not be fetched from this session's network and is deliberately absent from the table. The kernel numbers are defaults; both the period and the slice are tunable and some platforms tune them.

06

The evidence wall

Every source behind this page, graded. The network allowlist for this research session reached GitHub, GitLab, Google Cloud and Datadog only; the gaps that creates are listed at the end of the section.

Postmortem PostHog2025-10

Feature flags recurring outages

Four incidents in ten days; three shared one root cause, "nodes too small relative to pod resource requests", packing too many pods per node past 90 percent CPU. The fix raised requests and shrank the fleet.

Carry forwardDensity is not free efficiency; requests must cover the startup peak, and CPU saturation will masquerade as dependency failures.
github.com/PostHog/posthog.com · post-mortems
Postmortem PostHog2025-09

Feature flags service outage

A timeout cut from 1 s to 300 ms during database saturation; crash loops, retry amplification without circuit breakers, and probes that kept routing to dead pods for 45 minutes. 78 percent of US evaluations failed for 108 minutes.

Carry forwardTimeouts, retries and capacity margins are one budget; tune them against the loaded case, not the happy one.
github.com/PostHog/posthog.com · post-mortems
Postmortem GitLab2026-05

Incident review INC-10255: Redis primary CPU saturation

An audit-event surge saturated one single-threaded Redis primary; deferral created a re-enqueue feedback loop; audit events dropped globally for just over three hours.

Carry forwardUtilization planning is per-bottleneck. The fleet average says nothing about the component that cannot use a second core.
gitlab.com · gl-infra/production #22170
Source Cloudflare → Kubernetes2018-08

Issue #67577: CFS quotas can lead to unnecessary throttling

Opened by Ivan Babrou with a reproducer: quota throttles "well behaved tenants" that never consume their budget. Became the community's clearing-house for the limits argument for five years.

Carry forwardThe highest-signal document on this topic is an issue thread, not a blog post; read it before setting a limits policy.
github.com/kubernetes/kubernetes #67577
Source Linux kernel2018 (v4.18)

512ac999: fix bandwidth timer clock drift condition

The drift check "can never hit", so quota expiration silently never ran. Fixing the bug made enforcement honest, and honest enforcement exposed the throttling everyone then reported.

Carry forwardA correctness fix can be a performance regression for everyone who had calibrated against the broken behaviour.
github.com/torvalds/linux · 512ac999
Source Linux kernel (Indeed)2019 (v5.4)

de53fd7aed: remove expiration of cpu-local slices

Dave Chiluk's fix: threaded, non-CPU-bound apps were "throttled while simultaneously not consuming the allocated amount of quota" because up to 1 ms per core stranded every 100 ms period.

Carry forwardOn high-core-count nodes, small limits on threaded runtimes are structurally hostile; the stranding scaled with cores, not with usage.
github.com/torvalds/linux · de53fd7aed
Source Kubernetes2018-11

Issue #70585: disable cpu quota for Guaranteed pods

The original ask, framed by its reporter as a one-line conditional. Stayed open through two rejected pull requests and closed only when #127525 merged in 2025.

Carry forwardDate the ask, not the fix: whatever policy you set today on pinned pods was contested for six years.
github.com/kubernetes/kubernetes #70585
Source Kubernetes2019-11

PR #75682, closed unmerged: unset quota when CPU sets are in use

Died on a real objection: quota applied at the pod-level cgroup too, so the container-level exemption did not stop the throttling. The author closed it in favour of the kernel fix.

Carry forwardA rejected PR records why the obvious approach fails; this one says the enforcement hierarchy, not the flag, is what you must understand.
github.com/kubernetes/kubernetes #75682
Source Kubernetes2023-03

PR #107589, closed unmerged: no CPU quota for guaranteed pods

The second attempt, stalled on CRI side-effects across containerd and cri-o and the breadth of SIG review required, then closed for inactivity and revived by a third author.

Carry forwardChanges to enforcement semantics need every runtime and every constituency to agree; budget years, not sprints.
github.com/kubernetes/kubernetes #107589
Source Kubernetes2025-02

PR #127525, merged: no cfs quota for static-placement pods

The 2018 ask lands in v1.33 behind DisableCPUQuotaWithExclusiveCPUs, default on, with an explicit escape hatch for "users observing a regression".

Carry forwardFrom v1.33, pinned Guaranteed pods stop being throttled; below v1.33, pinning does not protect you from the quota.
github.com/kubernetes/kubernetes #127525
Source GitLab2020-09

gl-infra 11424: "CPU limits in kubernetes"

GitLab's platform team re-derives the question from the kernel bug, checks its GKE kernels against the fix, and asks the honest question: is it now safe to use CPU limits?

Carry forwardThe right form of the limits debate is kernel-version-specific, not ideological; GitLab's thread is the template.
gitlab.com · gl-infra #11424
Source GitLabchecked 2026-10

gitlab chart: webservice values.yaml shipped defaults

The product GitLab ships to everyone else defaults to a 300m CPU request with the limits block present and commented out; the sidekiq chart repeats the pattern.

Carry forwardWatch what operators ship, not what they debate: the commented-out block is the decision, recorded in the default.
gitlab.com · charts/gitlab webservice
Source Google / UMass2021

cluster-resource-forecast: the overcommit simulator

The published artifact of the EuroSys 2021 peak-prediction work: oversubscribe a machine by predicting its real peak from trace history instead of trusting summed requests. Built directly on the 2019 Borg traces.

Carry forwardOvercommit is a prediction problem; the safe level is discovered from usage history, not set as a fleet-wide percentage.
github.com/googleinterns/cluster-resource-forecast
Decision record Google / Kubernetes~2015

Resource QoS design: the overcommit contract

The founding document: requests are guaranteed, the request-to-limit gap is scavengeable, nodes are deliberately oversubscribable, and Borg's 20 percent gain is the stated justification. CPU is throttled, never killed; the ladder decides memory.

Carry forwardThe slack in your cluster is a designed feature with an intended consumer; if nothing scavenges it, you are running half the design.
github.com · design-proposals-archive · resource-qos.md
Decision record Kubernetes SIG Autoscaling2017+

Vertical Pod Autoscaler design proposal

Introduces itself as "aka rightsizing or autopilot": consume usage histograms and OOM events, recommend requests, apply them. The design's dependency on eviction is what the adoption number later indicted.

Carry forwardA rightsizer is only as adoptable as its apply path; recommendation quality was never the bottleneck.
github.com · design-proposals-archive · vertical-pod-autoscaler.md
Decision record Kubernetes SIG Nodestable v1.35

KEP-1287: in-place update of pod resources

Names the blocker in its motivation: resources were immutable, so every resize was a recreation, "a serious disruption" for stateful and low-replica workloads. Now stable, which removes the main recorded excuse for not rightsizing.

Carry forwardRe-evaluate any 2021-era "VPA is not worth it" decision; the constraint it was based on expired in v1.35.
github.com · enhancements · KEP-1287
Decision record AWS2022

Karpenter consolidation design

The repack loop, priced: delete a node when its pods fit elsewhere, replace it with a cheaper one otherwise, score candidates by a per-pod disruption cost, and make "minimal changes to the cluster" per step.

Carry forwardSavings equal the price of deleted nodes; any rightsizing effort that does not end in node deletion is accounting, not saving.
github.com/aws/karpenter-provider-aws · consolidation.md
Decision record Zalando~2016-2019

kubernetes-on-aws: container resource limits notes

A platform team's working notes, honest about their own uncertainty: "effect of CPU limits is not completely straight forward to understand", and higher limits than requests "allow to over-provision nodes, but has the danger of over-utilizing it".

Carry forwardEven platform teams running thousands of nodes documented CPU limit semantics as educated guessing; verify behaviour on your kernel, not from folklore.
github.com/zalando-incubator/kubernetes-on-aws · resource-limits.rst
Case study Datadog2023-11

Container Report: the measured size of the gap

Across Datadog's customer base: over 65 percent of Kubernetes workloads use less than half their requested CPU and memory; VPA adoption below 1 percent and flat since 2021; HPA above 50 percent.

Carry forwardThe gap is the norm, not your team's failure; and the tool built to close it is the least adopted tool in the ecosystem.
datadoghq.com/container-report
Vendor Google Cloud2018-05

Kubernetes best practices: resource requests and limits

The vendor framing of the contract: requests are guaranteed, "Kubernetes will make sure your containers get the CPU they requested and will throttle the rest", and the overcommitted state is presented as normal operation.

Carry forwardThe 2018 guidance assumes limits everywhere and predates both kernel fixes; date any best-practice post before applying it.
cloud.google.com · GKE best practices, 2018
Vendor Google Cloud2021-02

Introducing GKE Autopilot

The billing boundary moves: "you pay only for the pods you use and you're billed per second for vCPU, memory and disk resource requests. No more worries about unused capacity!" The unused capacity still exists; Google now owns it.

Carry forwardPer-request billing converts your inflated requests from invisible fleet slack into a visible line item; it is a rightsizing forcing function you pay for.
cloud.google.com · GKE Autopilot, 2021
Vendor Google2019-2020

ClusterData 2019: eight Borg cells of requests vs usage

Google's own framing of its public trace release: "the 2019 traces focus on resource requests and usage", with 5-minute CPU usage histograms rather than point samples, across eight cells for May 2019.

Carry forwardThe operator with the most data chose request-vs-usage as the thing worth publishing; the histograms exist because point samples hide the peaks that matter.
github.com/google/cluster-data · ClusterData2019.md
Vendor Kubernetes projectchecked 2026-10

Official docs: how limits are actually enforced

The enforcement asymmetry in the project's own words: CPU limits are "a hard limit the kernel enforces" by throttling; memory limits are enforced "reactively" by OOM kills that may lag the overage.

Carry forwardCPU and memory ceilings are different instruments; any policy that treats "limits" as one knob is wrong about at least one resource.
github.com/kubernetes/website · manage-resources-containers.md
What this wall is missing, and why

This session's network could not reach the canonical engineering-blog corpus (Uber's cpuset migration, Omio's throttling outage account, Indeed's two-part "Unthrottled", Buffer), the CAST AI utilization reports, any conference talk, or the paper PDFs for Borg (EuroSys 2015), Autopilot (EuroSys 2020) and the peak-prediction overcommit work (EuroSys 2021). Nothing from those sources is quoted here; where their claims matter, this guide cites the kernel commits, issue threads and artifact repositories that carry the same facts on reachable hosts. The widely circulated Autopilot slack figures are deliberately absent because they could not be verified against the paper. Those documents exist and are worth your time; the hunt queries in section 8 will find them from an unrestricted network.

07

Build a miniature, then productionise it

Six rungs from reproducing the throttle to defending a utilization target with evidence.

Reproduce throttling below the budget

On any Linux box with cgroups, run a multi-threaded, bursty workload under a CPU quota (a Java or Go service doing periodic fan-out is ideal). Watch cpu.stat: nr_periods, nr_throttled, throttled_time, against actual CPU usage.

Done when: you can show throttled periods while total usage sits well under the quota, and explain it with the 100 ms window.  Teaches: the ceiling is a latency instrument, not a capacity instrument.

Measure your own gap

For one real cluster, export requested vs used CPU and memory per workload over two weeks. Produce the two numbers that matter: total slack in cores, and the percentage of workloads under half their request (your private Datadog fact 5).

Done when: you can state "we reserve X cores and use Y at p95" with the query saved.  Teaches: the gap is measurable in an afternoon; the hard part was never the measurement.

Remove limits on one service, deliberately

Pick a latency-sensitive, trusted service. Drop its CPU limit, keep its request and memory limit, and compare p99 and throttled-period counts for a week against a control. Check the node's other tenants for interference.

Done when: you have before/after p99 and nr_throttled, plus a neighbour-interference check, written down.  Teaches: what the ceiling was costing you, and what (if anything) it was protecting.

Rightsize against the boot peak

Resize one deployment's requests from p95 steady-state usage plus its measured startup burst (time the pool initialisation and JIT warm-up under load). Then roll it and verify cold start still fits on a packed node.

Done when: a rolling restart under production load completes without startup-timeout crash loops.  Teaches: PostHog's lesson at miniature scale; the first minute is the sizing constraint.

Convert slack into deleted nodes

Enable a consolidation mechanism (Karpenter consolidation, or a manual drain-and- delete pass using the same arithmetic) and track node-hours per week before and after your rightsizing from rung 4.

Done when: the node count drops and the bill line follows; if requests shrank but nodes did not, you have proven the repack loop was the missing step.  Teaches: where savings actually materialise.

Run the saturation game day

On a staging cluster, pack a node past 85 percent CPU and roll a deployment onto it. Verify your alerts fire (PostHog settled on 80 percent sustained five minutes), your probes evict crash-looping pods quickly, and your retry budgets do not amplify.

Done when: the on-call is paged by CPU pressure before error rate, and the rollout degrades gracefully instead of cascading.  Teaches: whether your margins are insurance or decoration.

08

Keep hunting

The queries that found this material, grouped by what they surface. The vocabulary is the asset: throttling, CFS quota, bin-packing, slack, overcommit, rightsizing.

The enforcement record (kernel and Kubernetes)

  • "CFS quota" throttling kernel bug kubernetes unthrottled
  • repo:kubernetes/kubernetes is:issue "cfs quotas" throttling
  • repo:kubernetes/kubernetes is:pr is:closed is:unmerged cfs quota
  • repo:torvalds/linux commit "cfs_quota" throttling expiration

Incidents and operator decisions

  • postmortem "cpu throttling" OR "noisy neighbor" production root cause
  • site:github.com handbook post-mortems "resource requests"
  • gitlab.com gl-infra "incident review" CPU saturation
  • "cpu limits" site:gitlab.com gl-infra kubernetes

The economics

  • kubernetes cluster utilization average CPU report waste requests
  • container report "requested" utilization datadog OR "cast ai"
  • kubernetes "slack" requests usage "rightsizing" adoption

The research chain (paper → artifact → production)

  • overcommitment datacenter "peak prediction" oversubscription
  • google cluster-data borg traces requests usage histograms
  • autopilot workload autoscaling google eurosys slack
  • "avoiding cpu throttling" engineering cpuset containerized
09

References

  1. kubernetes/kubernetes #67577: CFS quotas can lead to unnecessary throttling GitHub, opened 2018-08-20 by bobrik. Checked 2026-10-03.
  2. Linux commit 512ac999: sched/fair: Fix bandwidth timer clock drift condition torvalds/linux mirror, merged for v4.18 (2018). Checked 2026-10-03.
  3. Linux commit de53fd7aed: sched/fair: Fix low cpu usage with high throttling (Dave Chiluk) torvalds/linux mirror, merged for v5.4 (2019). Checked 2026-10-03.
  4. kubernetes/kubernetes PR #75682: Unset CPU CFS quota when CPU sets are in use (closed unmerged) GitHub, closed 2019-11-09. Checked 2026-10-03.
  5. kubernetes/kubernetes PR #107589: kubelet: do not set CPU quota for guaranteed pods (closed unmerged) GitHub, closed 2023-03-31. Checked 2026-10-03.
  6. kubernetes/kubernetes #70585: Disable cpu quota (use only cpuset) for pod Guaranteed GitHub, opened 2018-11-02. Checked 2026-10-03.
  7. kubernetes/kubernetes PR #127525: static-placement pods should not have cfs quota enforcement GitHub, merged 2025-02-12, Kubernetes v1.33. Checked 2026-10-03.
  8. Resource Quality of Service in Kubernetes (design proposal) kubernetes/design-proposals-archive, ~2015. Checked 2026-10-03.
  9. Vertical Pod Autoscaler (design proposal) kubernetes/design-proposals-archive, 2017 onward. Checked 2026-10-03.
  10. KEP-1287: In-place update of pod resources kubernetes/enhancements; kep.yaml reads stage stable, latest-milestone v1.35. Checked 2026-10-03.
  11. Karpenter: Cluster consolidation design aws/karpenter-provider-aws, 2022. Checked 2026-10-03.
  12. Zalando kubernetes-on-aws: Container resource limits GitHub, ~2016-2019. Checked 2026-10-03.
  13. PostHog post-mortem: Feature flags recurring outages PostHog handbook (source file), 2025-10-21. Checked 2026-10-03.
  14. PostHog post-mortem: Feature flags service outage PostHog handbook (source file), 2025-09-29. Checked 2026-10-03.
  15. GitLab Incident Review INC-10255: Redis primary CPU saturation on redis-sidekiq nodes GitLab gl-infra/production, 2026-05-20. Checked 2026-10-03.
  16. GitLab gl-infra 11424: CPU limits in kubernetes GitLab, 2020-09-21. Checked 2026-10-03.
  17. GitLab chart: webservice values.yaml (shipped resource defaults) GitLab, master branch. Checked 2026-10-03.
  18. Datadog: 2023 Container Report Datadog, November 2023. Checked 2026-10-03.
  19. Google Cloud: Kubernetes best practices: Resource requests and limits Google Cloud blog, 2018-05-11. Checked 2026-10-03.
  20. Google Cloud: Introducing GKE Autopilot Google Cloud blog, 2021-02-25. Checked 2026-10-03.
  21. Google cluster-data: ClusterData 2019 traces GitHub, 2019-2020. Checked 2026-10-03.
  22. cluster-resource-forecast: Overcommit Simulator (EuroSys 2021 artifact) GitHub, 2021. Checked 2026-10-03.
  23. Kubernetes docs source: Resource Management for Pods and Containers kubernetes/website, main branch. Checked 2026-10-03.