Cost & Efficiency 03 Oct 2026 29 min read

The compute you reserve but never use

How production platforms close the gap between reserved and used compute: requests and limits, CFS throttling, overcommit, rightsizing, consolidation, and who pays for the slack.

Most containers use less than half of what they reserve, and the bill follows the reservation. This guide reconstructs the request/limit contract and its failure record from the primary sources this topic actually lives in: two Linux kernel commits, the six-year kubernetes quota-exemption argument (one issue, two rejected PRs, one merged fix), Google's QoS and VPA design records, Karpenter's consolidation pricing, GitLab's shipped chart defaults and incident reviews, PostHog's October 2025 postmortem, and Datadog's measured utilization figures. A reader leaves able to decide limits policy per kernel version, size requests against the boot peak, and say which loop actually turns slack into savings.

The finding that surprised me

The reservation number is wrong in both directions at once: Datadog measures over 65% of containers using less than half their requested CPU, while the freshest postmortem in the corpus (PostHog, October 2025) attributes three outages to requests that were too small and fixes them by raising requests, because a request is simultaneously the packing input, the performance insurance and the billing line, and those three jobs want different numbers.

What you get out of it

  • The bill follows summed requests, not usage; shrinking requests saves nothing until a repacking loop deletes nodes, which is why Karpenter prices savings as 'the price of the deleted node'.
  • The slack is a designed feature, not an accident: the Kubernetes QoS design oversubscribes nodes on purpose, citing Borg's ~20% utilization gain from scheduling into the request-to-limit gap.
  • CPU limits throttled workloads that never used their budget for years because of two kernel bugs (fixed v4.18 and v5.4); nr_throttled over nr_periods is the signal, and average-CPU dashboards cannot see any failure mode in this corpus.
  • Operators' real decisions live in defaults and issue threads, not blog posts: GitLab ships its chart with the CPU limits block commented out, and the pinned-pod quota exemption took one issue, two rejected PRs and 6.3 years to merge (v1.33, still gated).
  • Requests must cover the startup peak, not the steady state: PostHog's packed nodes passed 90% CPU, pool initialisation missed the 20-second startup timeout, and the cascade cost over 14 hours across ten days.

Scope

Why this, now. In-place pod resize went stable (v1.35) and the v1.33 quota exemption for pinned pods closed a six-year argument, which together expire the two best-documented excuses for the measured status quo of clusters reserving more than twice what they use.

What it does not cover. GPU fleets, serverless tiers, VM-level rightsizing, reservation and savings-plan financial engineering, and interruptible capacity (its own sibling dig); also, by network constraint rather than choice, the classic engineering-blog corpus (Uber, Omio, Indeed, Buffer), CAST AI's utilization reports, all conference talks and the Borg/Autopilot/peak-prediction paper PDFs, which are named as gaps in the ledger instead of cited.

Open the field guide → Self-contained: it loads nothing at read time, follows your system theme, and prints cleanly.