Evidence ledger
One row per claim in The compute you reserve but never use: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Topic: how production platforms close (or fail to close) the gap between the compute a workload reserves and the compute it uses: requests and limits, CFS throttling, overcommit, rightsizing, consolidation, and who pays for the slack.
All links checked 2026-10-03. Network note: this session's egress allowlist reached github.com, raw.githubusercontent.com, gitlab.com, cloud.google.com and www.datadoghq.com only. The classic engineering-blog corpus on this topic (Uber's cpuset migration, Omio's throttling account, Buffer, Indeed's "Unthrottled" pair), the CAST AI utilization reports, USENIX/ACM paper PDFs (Borg EuroSys'15, Autopilot EuroSys'20, "Take it to the limit" EuroSys'21, Borg-next-generation EuroSys'20) and all conference talks (Chiluk's KubeCon 2019 talk, Henning Jacobs' slide decks) were unreachable and are therefore named in the guide as evidence gaps rather than cited. Where a paper's claims matter, the guide cites the paper's own artifact repository on a reachable host instead.
| # | Org | Title | Tier | Published | Checked | URL | Claim I take from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | Cloudflare → Kubernetes | Issue #67577: CFS quotas can lead to unnecessary throttling | source | 2018-08-20 | 2026-10-03 | https://github.com/kubernetes/kubernetes/issues/67577 | Kubernetes CPU limits throttled workloads that were not consuming their quota; the thread became the community's clearing-house for the problem | Opened by bobrik (Ivan Babrou): CFS quotas "can lead to unnecessary throttling, especially for well behaved tenants"; body links kernel bugzilla 198197 and the kernel patch thread; reporter notes it is "not a bug in Kubernetes per se" |
| 2 | Linux kernel (Alibaba author) | commit 512ac999: sched/fair: Fix bandwidth timer clock drift condition | source | 2018 (v4.18) | 2026-10-03 | https://github.com/torvalds/linux/commit/512ac999d2755d2b7109e996a76b6fb8b888631d | The pre-4.18 clock-drift bug meant quota expiration never fired; fixing it exposed throttling everywhere | "The current condition to judge clock drift in expire_cfs_rq_runtime() is wrong, the two runtime_expires are actually the same when clock drift happens, so this condtion can never hit" |
| 3 | Linux kernel (Dave Chiluk, Indeed) | commit de53fd7aed: sched/fair: Fix low cpu usage with high throttling by removing expiration of cpu-local slices | source | 2019 (v5.4) | 2026-10-03 | https://github.com/torvalds/linux/commit/de53fd7aedb100f03e5d2231cfce0e4993282425 | Threaded, non-CPU-bound apps were throttled without consuming their quota because per-CPU slices stranded ~1ms per core per period | "highly-threaded, non-cpu-bound applications running under cpu.cfs_quota_us constraints can hit a high percentage of periods throttled while simultaneously not consuming the allocated amount of quota"; "typically no more than min_cfs_rq_runtime=1ms per cpu" |
| 4 | Kubernetes | PR #75682: Unset CPU CFS quota when CPU sets are in use (closed unmerged) | source | 2019 (closed 2019-11-09) | 2026-10-03 | https://github.com/kubernetes/kubernetes/pull/75682 | First attempt to exempt pinned pods from quota died on pod-level cgroup enforcement and the kernel fix | Author praseodym closing: "Proper implementation of this feature (or bug workaround, rather) will get pretty complicated, so I'm fully supportive of closing this PR in favour of the kernel fix"; derekwaynecarr had flagged that quota still applies at the pod-level cgroup |
| 5 | Kubernetes | PR #107589: kubelet: do not set CPU quota for guaranteed pods (closed unmerged) | source | 2022-01-16 (closed 2023-03-31) | 2026-10-03 | https://github.com/kubernetes/kubernetes/pull/107589 | Second attempt stalled on SIG review breadth and CRI side-effects, then died of inactivity | odinuge asked for broader SIG Node review citing CRI implementation impacts; closed after "please rebase" went unanswered; MarSik: "I posted the revive as #117030" |
| 6 | Kubernetes | Issue #70585: Disable cpu quota (use only cpuset) for pod Guaranteed | source | 2018-11-02 | 2026-10-03 | https://github.com/kubernetes/kubernetes/issues/70585 | The ask to stop throttling pinned Guaranteed pods was filed in 2018 and stayed open across both rejected PRs | "why not make a if statement when we add cpuQuota config, if it's guaranteed case, just set cpuQuota=-1?"; closed with resolution PR #127525 |
| 7 | Kubernetes | PR #127525: pods meeting qualifications for static placement ... should not have cfs quota enforcement | source | merged 2025-02-12 (v1.33) | 2026-10-03 | https://github.com/kubernetes/kubernetes/pull/127525 | The 2018 ask finally merged in 2025 behind a default-on feature gate | "containers meeting the qualifications for static cpu assignment...will not have cfs quota enforced"; gate DisableCPUQuotaWithExclusiveCPUs, "default on" |
| 8 | Google / Kubernetes | Design proposal: Resource Quality of Service in Kubernetes | adr | ~2015-2016 | 2026-10-03 | https://github.com/kubernetes/design-proposals-archive/blob/main/node/resource-qos.md | Kubernetes' request/limit split was designed to import Borg's overcommit gains; the QoS ladder decides who suffers when the gap is claimed | "This allows Kubernetes to oversubscribe nodes, which increases utilization, while at the same time maintaining resource guarantees for the containers that need guarantees"; "Borg increased utilization by about 20% when it started allowing use of such non-guaranteed resources, and we hope to see similar improvements in Kubernetes"; "Pods will not be killed if CPU guarantees cannot be met ... they will be temporarily throttled" |
| 9 | Kubernetes SIG Autoscaling | Vertical Pod Autoscaler design proposal | adr | 2017 onward | 2026-10-03 | https://github.com/kubernetes/design-proposals-archive/blob/main/autoscaling/vertical-pod-autoscaler.md | The community's rightsizing mechanism is named after Google's; its goal is utilization with bounded OOM/starvation risk | VPA "(aka. \"rightsizing\" or \"autopilot\") is an infrastructure service that automatically sets resource requirements of Pods"; goal 2: "Improving utilization of cluster resources, while minimizing the risk of containers running out of memory or getting CPU starved" |
| 10 | Kubernetes SIG Node | KEP-1287: In-place update of pod resources | adr | 2019, stable v1.35 | 2026-10-03 | https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/1287-in-place-update-pod-resources | Until 1.35, every resize was an eviction, which is why nobody resized; kep.yaml now reads stage: stable | "changing resource allocation requires the Pod to be recreated since the PodSpec's Container Resources is immutable"; "for stateful or batch workloads, Pod restart is a serious disruption"; kep.yaml: status: implemented, stage: "stable", latest-milestone: "v1.35" |
| 11 | AWS | Karpenter consolidation design | adr | 2022 | 2026-10-03 | https://github.com/aws/karpenter-provider-aws/blob/main/designs/consolidation.md | Reclaimed request-space becomes money only when a repacker deletes or replaces nodes; the design prices disruption against savings | "the cost savings is the price of the deleted node"; node replacement: "launch the new cheaper node and when it is ready delete the existing more expensive node"; "We intend to limit consolidation to making minimal changes to the cluster as we work towards reducing cluster cost" |
| 12 | Zalando | kubernetes-on-aws user guide: Container resource limits | adr | ~2016-2019 | 2026-10-03 | https://github.com/zalando-incubator/kubernetes-on-aws/blob/dev/docs/user-guide/resource-limits.rst | A platform team's own working notes state the overcommit trade plainly | "choosing higher limits than requests allows to over-provision nodes, but has the danger of over-utilizing it"; "requests are for making scheduling decisions"; "effect of CPU limits is not completely straight forward to understand" |
| 13 | PostHog | Post-mortem: Feature flags recurring outages (Oct 21-30, 2025) | postmortem | 2025-10-21 | 2026-10-03 | https://github.com/PostHog/posthog.com/blob/master/contents/handbook/company/post-mortems/2025-10-21-feature-flags-recurring-outages.md | Packing too many pods per node saturated CPU and caused three of four incidents; the fix was to raise requests, not shrink them | "Our nodes were too small relative to pod resource requests, causing Kubernetes to pack too many pods per node and saturate CPU capacity"; "over 14 hours of cumulative major impact"; node CPU "exceeding 90%"; fix: "Running smaller fleets with better-resourced pods rather than larger fleets with CPU-bound pods"; "CPU alerting was completely absent" |
| 14 | PostHog | Post-mortem: Feature flags service outage (Sep 29, 2025) | postmortem | 2025-09-29 | 2026-10-03 | https://github.com/PostHog/posthog.com/blob/master/contents/handbook/company/post-mortems/2025-09-29-flags-is-down.md | A timeout shrunk below real pool-initialization time under load converted saturation into a 108-minute outage | "approximately 78% of flag evaluation requests in the US region failed"; timeout cut "from 1 second to 300 milliseconds"; "Kubernetes continued routing traffic to pods in crash loops for up to 45 minutes" |
| 15 | GitLab | Incident Review INC-10255: Redis primary CPU saturation on redis-sidekiq nodes | postmortem | 2026-05-20 | 2026-10-03 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/22170 | Fleet-level slack does not protect a single-threaded component; its one core is the capacity that matters | "Redis's single-threaded architecture limited CPU scaling, compounding the resource bottleneck"; audit events dropped globally for ~3 hours; deferral "appears to have created a feedback loop" |
| 16 | GitLab | gl-infra issue 11424: CPU limits in kubernetes | source | 2020-09-21 | 2026-10-03 | https://gitlab.com/gitlab-com/gl-infra/production-engineering/-/work_items/11424 | GitLab's platform team re-derived the limits question from the kernel bug, kernel version by kernel version | "there was a nasty bug in the Linux scheduler resulting in throttling/performance issues of all processes, even those not reaching their cgroup limits"; "Both gstg and gprd GKE clusters are running on CoreOS version that's using 4.19 ... our GKE clusters should be free of this issue"; "Is it now (that the bug if fixed) safe to use CPU limits?" |
| 17 | GitLab | gitlab chart, webservice values.yaml (shipped defaults) | source | current (checked at master) | 2026-10-03 | https://gitlab.com/gitlab-org/charts/gitlab/-/blob/master/charts/gitlab/charts/webservice/values.yaml | GitLab ships its own product with CPU requests and no CPU limits; the limits block is present but commented out | resources: # limits: # cpu: 1.5 # memory: 3G requests: cpu: 300m memory: 2.5G (limits commented out in shipped default); sidekiq chart identical pattern |
| 18 | Datadog | 2023 Container Report (fact 5) | casestudy | 2023-11 | 2026-10-03 | https://www.datadoghq.com/container-report/ | Measured across Datadog's customer base: most containers use less than half of what they reserve, and almost nobody runs the rightsizer | "over 65 percent of Kubernetes workloads are utilizing less than half of their requested CPU and memory"; VPA adoption "below 1%", flat since 2021; "over 50% of Kubernetes organizations" run HPA |
| 19 | Google Cloud | Kubernetes best practices: Resource requests and limits (Sandeep Dinesh) | vendor | 2018-05-11 | 2026-10-03 | https://cloud.google.com/blog/products/containers-kubernetes/kubernetes-best-practices-resource-requests-and-limits | The vendor's own framing of overcommit: requests are guaranteed, the rest is throttled on contention | "Kubernetes goes into something called an 'overcommitted state' ... Because CPU can be compressed, Kubernetes will make sure your containers get the CPU they requested and will throttle the rest" |
| 20 | Google Cloud | Introducing GKE Autopilot | vendor | 2021-02-25 | 2026-10-03 | https://cloud.google.com/blog/products/containers-kubernetes/introducing-gke-autopilot | Autopilot moves the bin-packing problem (and the stranded capacity) to the vendor by billing per pod request | "you pay only for the pods you use and you're billed per second for vCPU, memory and disk resource requests. No more worries about unused capacity!" |
| 21 | cluster-data: ClusterData 2019 traces (John Wilkes) | vendor | 2019-2020 | 2026-10-03 | https://github.com/google/cluster-data/blob/master/ClusterData2019.md | Google published a month of request-vs-usage data for eight Borg cells precisely because the gap is the research problem | "The 2019 traces focus on resource requests and usage"; "CPU usage information histograms for each 5 minute period, not just a point sample"; eight cells, May 2019, ~2.4TiB compressed | |
| 22 | Google / UMass (EuroSys'21 artifact) | cluster-resource-forecast: Overcommit Simulator | source | 2021 | 2026-10-03 | https://github.com/googleinterns/cluster-resource-forecast | The published artifact of the peak-prediction overcommit work: oversubscribe machines by predicting the peak, not by trusting requests | "mimics Google's production environment in which Borg operates"; implements "the logic to run peak-oracle and other practical predictors" |
| 23 | Kubernetes | Official docs source: Resource Management for Pods and Containers | vendor | current (checked at main) | 2026-10-03 | https://github.com/kubernetes/website/blob/main/content/en/docs/concepts/configuration/manage-resources-containers.md | The enforcement asymmetry in the project's own words: CPU is throttled, memory is killed | "cpu limits are enforced by CPU throttling ... a cpu limit is a hard limit the kernel enforces"; "memory limits are enforced by the kernel with out of memory (OOM) kills ... enforced reactively" |
Tier mix
postmortem 3 · source 10 · adr 5 · casestudy 1 · vendor 4 — 23 rows, 23 distinct documents, 5 hosts (github.com, raw.githubusercontent.com via blob pages, gitlab.com, cloud.google.com, www.datadoghq.com), ~12 organisations (Cloudflare, Linux kernel maintainers, Indeed via the kernel commit, Alibaba via the kernel commit, Kubernetes SIG Node and SIG Autoscaling, Google, Zalando, AWS, PostHog, GitLab, Datadog, UMass/Google).
Named gaps (unreachable this session, deliberately not cited)
- Borg (EuroSys'15), Autopilot (EuroSys'20), "Borg: the Next Generation" (EuroSys'20) and "Take it to the limit" (EuroSys'21) paper PDFs: ACM, USENIX and authors' sites all blocked. The Autopilot slack figures (23% vs 46%) circulate widely but could not be verified against the paper here, so the guide does not quote them.
- Uber "Avoiding CPU Throttling in a Containerized Environment" (2022), Omio "CPU limits and aggressive throttling in Kubernetes" (2020), Indeed "Unthrottled" parts 1-2 (2019), Buffer's account: hosts blocked. These are the canonical engineering-blog accounts; the same facts are carried here by the kernel commits, issue #67577 and the GitLab debate thread.
- CAST AI Kubernetes cost reports (2024-2026, headline ~8-13% average CPU utilization): host blocked; reported figures are mentioned in prose as unverified-this-session.
- All conference talks (Chiluk, KubeCon 2019; Jacobs, several 2018-2019): video and slide hosts blocked. Talk tier is empty in this guide for that reason.