The discount is a repossession clause
How production systems run on interruptible (spot / preemptible) cloud capacity: the 30-to-120-second repossession contract, the seven-part machinery that absorbs it, and the point where the 60-91% discount stops paying for that machinery.
Every large cloud sells the same machines twice: firm capacity at list price and the same capacity at 60-91% off with the right to take it back on 30 to 120 seconds' best-effort notice. This guide reconstructs, from Karpenter's design record, aws-node-termination-handler's issue history, KEP-2000, the three providers' own contract language, GitLab's production incident reviews and the NSDI '24 spot-traces artifact, what running on the discounted tier actually requires: a per-workload admission policy, ten-plus-type diversification, best-effort signal plumbing, a drain path that fits a 30-second window, a fallback ladder whose latency must be measured rather than assumed, paid headroom, and a churn budget. It closes with the decision conditions that flip each choice and a six-rung ladder ending in a capacity game day.
AWS's own best-practices documentation strongly warns against the spot-with-on-demand-failover pattern for capacity-intolerant workloads, the exact pattern the ecosystem's tooling implements as standard; and Karpenter's designers refused to ship pure price optimization because it would walk the fleet down the price ladder into the most-interrupted pools, so they gated it behind a 15-cheaper-types availability heuristic that users still report as a bug.
What you get out of it
- Capacity, not price, is what reclaims your instance: since bidding ended in 2017 spot prices move slowly, and AWS lists 'EC2 needs it back' first among interruption reasons; even the two-minute warning is emitted on a best-effort basis.
- The pool, not the discount tier, is the failure domain: GitLab's 2h55m runner outage and 77.2% scale-up failure rate happened on on-demand capacity, which is why an unrehearsed on-demand fallback is a hope rather than an availability design.
- Price and availability are one dial: Karpenter withheld spot-to-spot consolidation until it could proxy availability with a 15-cheaper-types flexibility floor, because providers do not expose pool capacity at all.
- The fallback path has minutes of unmeasured latency: cluster-autoscaler marked a dry spot group unhealthy only after ~10 minutes and wedged on placeholder instances; measure signal-to-serving under real scarcity or you do not know it.
- Just-in-time capacity converges back to paying for slack: balloon-pod workarounds became Karpenter's CapacityBuffer API, so the honest spot business case prices in headroom and churn, not just the discount.
Scope
Why this, now. Karpenter's capacity-buffers RFC and its gated spot-to-spot consolidation are both 2024-26 artifacts of the same lesson, that cost automation on reclaimable capacity needs an availability term, and GitLab's April-August 2026 stockout incidents supply fresh public evidence that the fallback tier itself runs dry.
What it does not cover. GPU training on spot (covered by the sibling dig on keeping training runs alive), serverless tiers, reservations and savings-plan financial engineering, and the classic engineering-blog corpus (Netflix, Delivery Hero, Honeycomb) plus AWS's own site and USENIX, which this session's network allowlist could not reach; those are named as gaps in the ledger rather than cited.
Other field guides
The meter runs on cardinality
Reconstructs the telemetry cost-containment funnel from the incident record at GitLab and Datadog, the pricing sheets of three clouds, Prometheus and…
37 sources · 20 organisations · 8 postmortemsThe bill is not a brake
Reconstructs the spend-containment architecture from six published billing incidents (Milkie Way, Troy Hunt, the Netlify and empty-S3-bucket bills, C…
34 sources · 31 organisations · 5 postmortemsWhen owning hardware wins, and when it owns you
Dropbox banked an SEC-audited $74.6M by leaving S3; 37signals cut its bill from $3.2M to $1.3M on the way to deleting its AWS account; and in the sam…
29 sources · 25 organisations · 3 postmortems