Capacity the cloud takes back  / field guide
Practitioner field guide · 2026-09-29

The discount is a repossession clause

Every large cloud sells the same machines twice: firm capacity at list price, and the same capacity at 60–91% off with the right to take it back on 30 to 120 seconds' notice. This guide reconstructs, from Karpenter's design record, three GitLab production incident reviews, the providers' own contract language and an NSDI '24 trace artifact, what it actually takes to run production on the discounted tier, and where the discount stops paying for the machinery that absorbs it.

27 primary sources 9 organisations 4 incident records Evidence through Sep 2026 Read: ~25 min
01

The territory

One physical fleet, two contracts, and the engineering that lives in the difference between them.

90–91%
Maximum claimed discount off on-demand: AWS "up to 90%", Google Cloud "up to 91%"
30–120s
Notice before repossession: 2 minutes on AWS, 30 seconds on GCP and Azure, all explicitly best-effort
15
Cheaper instance types Karpenter demands before it will consolidate one spot node onto another
2h 55m
CI runner shard outage at GitLab when one zone's VM pool ran dry, April 2026

State the problem without naming a product and it becomes clearer than the product pages make it. A cloud region is a warehouse of machines bought against forecast demand. Whatever the forecast missed sits idle, so the provider sells the idle fraction at a steep discount with a repossession clause attached: when a full-price customer shows up, the discounted tenant gets a short warning and then the machine back. AWS calls this Spot, Google and Azure call it Spot VMs, and all three write the clause in almost identical language. Google's product page is the most candid about the recovery half of the bargain: managed instance groups will recreate preempted VMs, and then, in parentheses, "if capacity is available".

The architecture question is therefore not "is the discount real". It is real, and since AWS abolished bidding in late 2017 the price side has been boring by design; the EC2 documentation describes prices as "adjusted gradually based on the long-term supply of and demand", and lists capacity, not price, as the reason instances are reclaimed: "Amazon EC2 can interrupt your Spot Instance when it needs it back". The question is what machinery your platform needs so that repossession is an accounting event rather than an incident, and whether that machinery, once built and operated, still costs less than the list price did.

This guide covers production service and batch capacity on interruptible VMs, mostly as seen through Kubernetes-shaped platforms, because that is where the public record is. It deliberately excludes GPU training on spot (a sibling dig, Keeping the training run alive, covers surviving preemption mid-training), serverless tiers, and financial engineering of reservations and savings plans. A network note bounds the evidence: this session could reach GitHub, GitLab, Google Cloud and Datadog, but not Netflix's, Delivery Hero's or Honeycomb's engineering blogs, AWS's own site, or USENIX. Named accounts from those sources are flagged as unfetched in the ledger and never cited.

Figure 1 · One fleet, two contracts

shortfall surfaces as
stockout on launch

shortfall surfaces as
repossession,
30-120s notice

fallback launches

One physical fleet per zone

Firm tier:
on-demand and reserved,
list price

Interruptible tier:
spot, 60-91% off

Your firm fleet

Your interruptible fleet

shortfall surfaces as
stockout on launch

shortfall surfaces as
repossession,
30-120s notice

fallback launches

One physical fleet per zone

Firm tier:
on-demand and reserved,
list price

Interruptible tier:
spot, 60-91% off

Your firm fleet

Your interruptible fleet

Both tiers draw on the same physical pool, so the fallback path from the interruptible tier lands on capacity that can itself be short; GitLab's stockout incidents happened entirely on the firm tier. Sources: AWS docs, GitLab INC-9343.
Diagram source
02

How it is actually built

Seven parts recur across every system in the record: an admission policy, pool diversification, signal plumbing, a vacate path, a replacement ladder, paid headroom, and churn observability.

Figure 2 · Reference architecture for interruptible capacity

On each node

Fleet shaping

Governance

Admission policy:
which workloads may ride spot,
with rollback metrics

Diversification:
10+ instance types,
every zone

Capacity-aware allocation
price-capacity-optimized

Paid headroom:
buffers or a static pool

Signal listener:
metadata poll or event queue

Cordon and drain
inside the notice window

Replacement launch
and fallback ladder

Churn observability:
interruption counts,
pod startup duration

On each node

Fleet shaping

Governance

Admission policy:
which workloads may ride spot,
with rollback metrics

Diversification:
10+ instance types,
every zone

Capacity-aware allocation
price-capacity-optimized

Paid headroom:
buffers or a static pool

Signal listener:
metadata poll or event queue

Cordon and drain
inside the notice window

Replacement launch
and fallback ladder

Churn observability:
interruption counts,
pod startup duration

Reconstructed from the aws-node-termination-handler, Karpenter's disruption docs, GKE's preemption handling and GitLab's per-workload change record. The admission policy and the headroom are the two parts teams add last and need first.
Diagram source

Admission is a policy decision, recorded per workload. The most concrete evidence of how a careful organisation adopts spot is a GitLab production change record from October 2024: one service (Mimir distributors, a stateless write path), one reviewed change to allow scheduling on spot, a named rollback metric ("an undesirable increase in Mimir write latency") and a five-minute revert plan. Datadog's Karpenter guidance states the same policy from the other side: workloads that cannot tolerate interruption should be explicitly limited to on-demand or reserved capacity. Nobody in the record adopts spot cluster-wide in one motion.

Diversification is the load-bearing wall, not a tuning tip. A spot "pool" is one instance type in one zone, and reclaim events are correlated inside a pool. AWS's own rule of thumb is to be "flexible across at least 10 instance types for each workload" with every zone enabled, and its allocation strategy (price-capacity-optimized) exists to spread launches toward the pools least likely to be reclaimed. Karpenter's designers went further and quantified the wall's minimum thickness: 15 cheaper candidate types before spot-to-spot consolidation is allowed, a number they derived from "analysis done on the flexibility of AWS customers request in the launch path". Below that flexibility, section 3 explains, the cost optimizer becomes a reliability hazard.

The signal plumbing has two shapes, and both are best-effort. The aws-node-termination-handler documents them: a DaemonSet on every node polling the instance metadata service, or a central deployment consuming EventBridge events from an SQS queue, which also carries Auto Scaling lifecycle hooks and health events. AWS recommends polling metadata every five seconds, and the same page carries the sentence that should size your ambitions: "Interruption notices are emitted on a best effort basis." The warning itself is not guaranteed. Choosing hibernation forfeits the two minutes entirely, because hibernation begins immediately.

The vacate path reshaped Kubernetes itself. Graceful node shutdown, KEP-2000, exists because externally triggered shutdowns, "e.g. Preemptible VMs" in the KEP's own motivation, killed pods without any termination lifecycle. The kubelet now takes a systemd inhibitor lock, delays the OS shutdown, and evicts pods in priority order within a configured grace budget. Two operational details matter: the feature defaults to a grace period of zero, so it does nothing until configured, and GKE only wires it up for you; Google's Spot VMs GA post describes the kubelet noticing the 30-second preemption notice and terminating pods gracefully as the managed behaviour. Your pods' terminationGracePeriodSeconds must fit inside the provider's notice, which on GCP and Azure is 30 seconds, not two minutes.

Replacement is where the tools diverge. Karpenter listens for the interruption warning and, per its disruption documentation, begins draining and provisioning a replacement at the same time, claiming that average node startup generally beats the two-minute clock. The cluster-autoscaler generation of tooling instead retries through prioritised node groups and discovers capacity problems by timeout, which section 4 shows can take tens of minutes. The third generation adds paid headroom back on top: Karpenter's capacity-buffers RFC formalises what users were already doing with "balloon pods, pause containers, or NodePools with fixed capacity", which is to say: teams running just-in-time capacity keep buying a slice of guaranteed slack, because the reclaim-and-replace loop is not fast enough on its own.

Figure 3 · A clean interruption, end to end

Replacement nodeDoomed nodeOrchestratorSignal listenerCloud control planeReplacement nodeDoomed nodeOrchestratorSignal listenerCloud control planepar[drain while replacing]interruption notice, T-120s, bestefforteventtaint and cordonevict pods, PDB-awarelaunch from healthiest cheap poolreadyreclaim at T-0, ready or not
Replacement nodeDoomed nodeOrchestratorSignal listenerCloud control planeReplacement nodeDoomed nodeOrchestratorSignal listenerCloud control planepar[drain while replacing]interruption notice, T-120s, bestefforteventtaint and cordonevict pods, PDB-awarelaunch from healthiest cheap poolreadyreclaim at T-0, ready or not
The drain and the replacement run in parallel because the notice window is too short to run them in sequence; the provider reclaims at T-0 whether or not either finished. Reconstructed from Karpenter's disruption docs and AWS's notice documentation.
Diagram source

Observability closes the loop on churn, not just downtime. The metric Datadog's Karpenter guide leads with is a plain counter of interruption events, and its reading advice is the architecturally interesting part: rising interruptions "can create brief capacity gaps", and a corresponding rise in pod startup duration "confirms that these interruptions aren't just transparent cost optimization". That pairing, reclaim rate against workload-visible latency, is the difference between knowing your discount's price and guessing it.

03

The decisions that matter

The public record contains three genuinely argued decisions: whether to act on the early-warning signal, whether to let the cost optimizer touch spot at all, and what the fallback tier is really worth.

Decision: act on the rebalance recommendation, or only on the real notice?

Chosen
  • Karpenter: publish the rebalance event, act only on the two-minute interruption notice. Its docs state it plainly: "Karpenter does not currently support taint, drain, and terminate logic for Spot Rebalance Recommendations" (disruption docs).
Rejected (but shipped elsewhere)
  • aws-node-termination-handler added drain-on-rebalance in 2021 after issue #368: "Two minutes may not sufficient time to drain node on some circumstance", for example when a replacement node must first be created.
Flips when
  • Your drain-plus-replace exceeds the notice window: act early despite churn.
  • Otherwise the signal is a poor trigger: AWS's own docs concede the rebalance recommendation "can arrive along with the two-minute interruption notice", so the early warning may not be early, and draining on an elevated-risk guess doubles your voluntary churn.

Decision: may the cost optimizer consolidate spot onto cheaper spot?

Chosen
  • Withheld for years, then shipped (v0.34, behind a feature gate) with a guard: single-node spot-to-spot consolidation requires at least 15 cheaper candidate instance types (design doc).
Rejected
  • Pure price minimization: "it would repeatedly interrupt the same instance, walking down the PCO decision ladder, until only the lowest price remained".
  • A minimum price-improvement factor (hard to tune), optimistic probe launches (billing confusion, riskier bugs), minimum node lifetime (converges to lowest price anyway). All three are recorded in the design doc as options 2–4.
Flips when
  • Consolidating many nodes into one: the guard is waived because a many-to-one move cannot race to the bottom.
  • Your NodePool allows fewer than ~15 types: the optimizer stays off and underutilised nodes idle, which users read as a bug (#1124, #1653: a t3.2xlarge at 26% CPU allocation left running).

Decision: what is the on-demand fallback actually for?

Chosen
  • Everyone builds one: cluster-autoscaler priority expander with an on-demand group last, Karpenter's capacity-type ordering (reserved, then spot, then on-demand, per Datadog's architecture guide).
The vendor's dissent
  • AWS's best-practices page warns that spot is "not recommended for workloads that are intolerant of occasional periods when the target capacity is not completely available", and then, remarkably: "We strongly warn against using Spot Instances for these workloads or attempting to fail-over to On-Demand Instances to handle interruptions" (spot-best-practices).
Flips when
  • The fallback is sized and rehearsed as surge capacity for a capacity-tolerant workload: fine.
  • The fallback is the availability story for a capacity-intolerant workload: it fails exactly when needed, because firm-tier launches draw on the same finite zone (section 4's stockouts) and the fallback path itself has minutes of latency (autoscaler #3490).

The second decision is the one worth internalising, because it generalises: on reclaimable capacity, price and availability are the same dial viewed from opposite sides. The Karpenter designers could not read availability directly, since, as the design doc notes, "cloud providers don't surface capacity pool information (availability) for spot", so they built a proxy for it out of flexibility and then refused to let the optimizer run without it. The complaints in #1124 and #1653 are the cost of that refusal made visible, and the maintainers labelled the second one "awaiting-more-evidence" rather than relaxing the guard. That is the correct order of operations, and most home-grown cost automation gets it backwards.

Figure 4 · Should this workload ride interruptible capacity?

no

yes

no

yes

no

yes

Survives losing any node
on 30-120s notice?

Firm capacity only.
Record the decision.

Flexible across 10+ types
and every zone?

Widen flexibility first;
spot under a narrow request
churns and stalls

Tolerates hours of reduced
capacity if pools run dry?

Firm baseline + paid headroom;
spot carries only the burst

Run on spot with
capacity-aware allocation
and churn budgets

no

yes

no

yes

no

yes

Survives losing any node
on 30-120s notice?

Firm capacity only.
Record the decision.

Flexible across 10+ types
and every zone?

Widen flexibility first;
spot under a narrow request
churns and stalls

Tolerates hours of reduced
capacity if pools run dry?

Firm baseline + paid headroom;
spot carries only the burst

Run on spot with
capacity-aware allocation
and churn budgets

The tree the sources collectively encode. The middle exit is the one most production services take: spot for the elastic fraction over a firm baseline. Derived from AWS guidance, Datadog's Karpenter guide and the capacity-buffers RFC.
Diagram source
DecisionChosenRejectedBecauseEvidence
Trigger signalTwo-minute notice (Karpenter)Rebalance recommendation (NTH drains on it, opt-in)Early warning is probabilistic and may arrive with the notice anywayKarpenter docs, NTH #368
Spot-to-spot consolidationGated on 15 cheaper typesPure price minimization; PIF; probe launches; min lifetime"Walking down the ladder" converts savings into interruptionsDesign doc
Signal transportCentral queue (SQS/EventBridge) for fleetsPer-node metadata polling as the only pathQueue mode also carries ASG lifecycle and health events; polling is per-node toilNTH README
Azure eviction policyDelete (disks go, no storage bill)Deallocate (default: keeps disks, counts against quota)Reallocation after deallocate is explicitly not guaranteedAzure docs, 2026
Price capNo cap (pay current spot price)Max-price biddingAWS: a max price makes interruptions "more frequently than if you do not specify it"AWS docs
HeadroomPaid buffers / static pools re-added on top of JITPure just-in-time provisioningBalloon-pod workarounds were universal enough to standardiseCapacity-buffers RFC
Adoption unitPer-workload allowlist with rollback metricCluster-wide default-to-spotInterruption tolerance is a property of a workload, not a clusterGitLab change 18749
04

What broke in production

The published incidents cluster into three classes: the pool was empty (at any tier), the fallback was slow, and the automation over-reacted. Notably absent from the public record: a written postmortem of a mass spot reclamation at a named company.

The class most people miss

GitLab's stockout incidents below happened on on-demand capacity. They are in this guide because they falsify the assumption the whole fallback pattern rests on: that the firm tier is always in stock. Spot users meet empty pools first and most often, but the pool, not the discount, is the failure domain. This is precisely why AWS warns against treating on-demand failover as an availability strategy for capacity-intolerant workloads.

Postmortem

The zone ran dry and took the CI fleet with it

AssumptionRunner-manager VMs can always be created; zonal capacity is effectively infinite for n2d-standard-2.
What happenedA GCP quota-calculation error in a storage layer surfaced as ZONE_RESOURCE_POOL_EXHAUSTED on every VM insert in us-east1-d. All six runner managers for the shard were pinned to that zone, so the shard could create no capacity at all.
Blast radius2 h 55 m; apdex from ~100% to ~20%; pending CI queue peaked around 24,000 jobs (GitLab, April 2026).
FixReconfigure runner managers to us-east1-c during the incident; follow-up to expand hosted runners to a third zone.
Design ruleA fallback zone you have never launched into is a hope, not a design. Diversification only counts if the launch path exercises it; pinning collapses it silently. Canary launches (GitLab used manual gcloud creates) are the ground truth when the provider's status page lags, which here it did by roughly three hours.
Postmortem

The machine type you standardised on stopped existing

AssumptionThe provider always has the machine family you standardised on; scale-up failures are your bug.
What happenedZonal shortages of the required machine types: GKE node scale-ups failed at rates up to 77.2% over 24 hours. Google confirmed "newly delivered machines were misallocated into a different product pool", a supply-chain error entirely outside the customer's control.
Blast radiusDelayed automatic scaling for a day (May 2026); a sibling record from August 2026 shows the same class stretching single deploy jobs to 45 minutes against N4/N4D limits and stockouts.
FixFallback node pools on other machine families; longer term, GitLab folded machine-family flexibility into its GCP contract negotiation, because per-SKU discounts fix your families for the contract's life.
Design ruleInstance-type flexibility is not only a spot tactic; it is leverage against provider-side supply errors, and it must be decided before the discount contract freezes your SKUs.
Issue thread

The fallback existed and still took tens of minutes

AssumptionPriority expander plus an on-demand group means spot shortage costs one max-node-provision-time (configured: 5 minutes), then capacity arrives.
What happenedWith GPU spot unavailable, cluster-autoscaler marked the spot group unhealthy only after ~10 minutes, created "placeholder" instances for it, and kept selecting it; the on-demand group was not scaled in the configured window.
Blast radiusPods pending for the full detection-plus-retry cycle per attempt (reported September 2020; the issue rotted closed without a recorded fix).
FixNone recorded in-thread. The ecosystem's structural answer became Karpenter-style flexible launches, where one request carries many acceptable types instead of retrying groups serially.
Design ruleMeasure your fallback's signal-to-serving latency under real scarcity, not its existence. A fallback that engages after the queue has built is a second incident, not a mitigation.
Design record

The optimizer that would have eaten its own fleet

AssumptionCheaper is better: a consolidation loop should always move workloads to the lowest-price node that fits.
What happenedA failure prevented by design rather than survived: Karpenter's designers worked out that pure price minimization on spot "would repeatedly interrupt the same instance, walking down the PCO decision ladder, until only the lowest price remained", concentrating the fleet in the most-contended pools. The mirror image shipped in NTH: draining on every rebalance recommendation, a signal AWS admits may arrive with the real notice, buys early warning at the price of doubled churn.
Blast radiusCounterfactual for Karpenter; for over-eager draining, user-visible as elevated node replacement and pod startup latency (the metric pair Datadog recommends watching).
FixAn availability heuristic (15 cheaper types) as a precondition for the optimizer; rebalance-draining left off by default in Karpenter.
Design ruleCost automation on reclaimable capacity needs an availability term in its objective, even a crude one, or it optimises the fleet into the pools most likely to vanish. Give every optimizer a churn budget and alert on the budget, not on individual interruptions.

Figure 5 · The stockout, as the runner managers saw it

GCE zoneus-east1-dRunner managers (allsix in us-east1-d)CI job queueGCE zoneus-east1-dRunner managers (allsix in us-east1-d)CI job queueloop[for 2h55m, queuegrows toward 24K]operators re-pinshard to us-east1-cjobs arriveinsert n2d-standard-2 VMZONE_RESOURCE_POOL_EXHA-USTEDmore jobsretry insertsexhaustedinsert VM in us-east1-ccreated, queue drains
GCE zoneus-east1-dRunner managers (allsix in us-east1-d)CI job queueGCE zoneus-east1-dRunner managers (allsix in us-east1-d)CI job queueloop[for 2h55m, queuegrows toward 24K]operators re-pinshard to us-east1-cjobs arriveinsert n2d-standard-2 VMZONE_RESOURCE_POOL_EXHA-USTEDmore jobsretry insertsexhaustedinsert VM in us-east1-ccreated, queue drains
Nothing was reclaimed and nothing crashed; capacity simply could not be created, and every retry landed in the same pinned zone until operators re-pointed it. From GitLab's INC-9343 review.
Diagram source

Figure 6 · Walking down the ladder: the loop Karpenter refused to ship

yes: consolidate

no: floor reached

Node in pool P:
healthy, price acceptable

Cheaper pool
available?

Interrupt workload,
relaunch in cheaper pool

Now nearer the bottom
of the price ladder

Higher contention:
reclaim more likely

Reclaimed; relaunch;
optimizer evaluates again

Fleet parked in the
most-contended pools

yes: consolidate

no: floor reached

Node in pool P:
healthy, price acceptable

Cheaper pool
available?

Interrupt workload,
relaunch in cheaper pool

Now nearer the bottom
of the price ladder

Higher contention:
reclaim more likely

Reclaimed; relaunch;
optimizer evaluates again

Fleet parked in the
most-contended pools

Each consolidation step is locally rational and the fixed point is a fleet concentrated in the cheapest, most-reclaimed pools. The 15-type guard breaks the loop by requiring broad headroom below the current price before any step is taken. From the spot consolidation design doc.
Diagram source

One absence deserves its own sentence. Between them, the sources here document stockouts, slow fallbacks and self-inflicted churn, but no fetched source in this corpus is a formal postmortem of a correlated mass spot reclamation taking down a named company's service. Practitioner accounts of that shape exist (several were surfaced by search against blocked hosts, and vendor material retells them anonymously), so read the gap narrowly: either teams absorb mass reclaims without customer impact, or the incidents happen and are not written up publicly. Both readings argue for rehearsing the event yourself rather than waiting for someone else's write-up.

05

Numbers you can plan against

The contract numbers are vendor-stated and stable; the operational numbers are single-organisation measurements. The one number nobody publishes is per-pool reclaim probability, which is why every design above builds a proxy for it.

MetricValueAtContextAs ofSource
Max claimed discountup to 90% / 91%AWS / GCPVendor claims; actual discount varies by pool and time2023 / 2026AWS, GCP
GA discount band60–91%GCPRange stated at Spot VMs general availability2022-05GCP blog
Interruption notice120 sAWSBest effort; poll metadata every 5 s; hibernation gets no lead time2023AWS docs
Eviction notice30 sGCP, AzureBoth best effort; Azure: "no SLA", "no high availability guarantees"2026GCP, Azure
Eviction-rate unit% per hourAzure"10% means a 10% chance of being evicted within the next hour", from 7-day history2026-02Azure docs
Flexibility floor (guidance)≥10 typesAWS"Good rule of thumb", plus all zones enabled2023AWS docs
Flexibility floor (enforced)15 cheaper typesKarpenterPrecondition for single-node spot-to-spot consolidation2024Design doc
Fallback engage time~10 min observedcluster-autoscaler userSpot group marked unhealthy at ~2× the configured 5-minute provision timeout2020-09Issue #3490
Stockout blast radius2 h 55 m / 24K jobsGitLabOne zone's pool exhausted; apdex ~100%→~20%2026-04INC-9343
Scale-up failure rateup to 77.2%GitLabMachine-type shortage, 24-hour window, on-demand tier2026-05INC-9866
Recurrence of the class2021, 2026×3GitLabSSD stockout 146 min (2021); zonal stockouts in Apr, May, Aug 20262021–2026#4940, #22668
Preemption measurability45-day tracesSkyPilot / UC BerkeleyAvailability and preemption traces on AWS and GCP, collected by pinging providers; artifact of an NSDI '24 Outstanding Paper2024spot-traces
Read these carefully

The discount ceilings are marketing numbers: they describe the best pool, not your blend. The operational numbers (2h55m, 77.2%, ~10 minutes) are honest but each comes from a single organisation and incident. Real per-pool interruption rates are published only as coarse bands (AWS's Spot Instance Advisor, unreachable from this session's network; Azure's portal eviction bands): treat any precise interruption-rate figure you are quoted as unverified, and measure your own, which rung 5 below makes cheap.

06

The evidence wall

Every source behind this page, graded. The mix is honest about its constraint: this session's network could reach GitHub, GitLab, Google Cloud and Datadog, so the classic engineering-blog accounts (Netflix's internal spot market, Delivery Hero, Honeycomb) are named in the ledger as unfetched rather than cited here.

Postmortem GitLab2026-04

Incident review INC-9343: CI runner shard SLO violation

A GCP zonal capacity failure surfaced as ZONE_RESOURCE_POOL_EXHAUSTED; all six runner managers were pinned to the failing zone. 2h55m, apdex to ~20%, 24K queued jobs. Operators diagnosed the zone ~3 hours before the provider's public acknowledgement.

Carry forwardZone diversification that the launch path does not exercise is decorative. Keep a canary launch as ground truth.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/21838
Postmortem GitLab2026-05

INC-9866: GKE node scale-up errors from machine-type shortage

Scale-ups failed at up to 77.2% over 24 hours because the provider had misallocated newly delivered machines between product pools. The customer's fix was family-level fallback node pools.

Carry forwardThe firm tier stocks out too; machine-family flexibility is an availability control, not just a spot tactic.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/21996
Postmortem GitLab2021-06

GKE unable to scale: SSD availability

An earlier instance of the same class: 146 minutes of impact when local-SSD-backed nodes could not be created in-zone. Same failure shape five years before INC-9343.

Carry forwardStockouts are a recurring class, not a freak event; design and drill for them on that basis.
gitlab.com/gitlab-com/gl-infra/production/-/issues/4940
Case study GitLab2026-08

N4 quota limits and N4D stockouts blocking deploys

Three consecutive production deploys slowed to 45 minutes by zonal stockouts; the investigation explicitly feeds machine-family strategy into a GCP contract negotiation before per-SKU discounts freeze.

Carry forwardDecide instance-family flexibility before the discount contract is signed, not after.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/22668
Case study GitLab2024-10

Change record: Mimir distributors allowed onto spot

How a mature operator actually adopts spot: one stateless service, a reviewed change, a named rollback metric (Mimir write latency) and a five-minute revert path.

Carry forwardAdopt per workload, with the rollback metric written down before the change ships.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/18749
Design record Karpenter2023/24

Spot consolidation design

The central ADR of this topic: why pure price minimization on spot self-destructs ("walking down the PCO decision ladder"), why availability must be proxied by a flexibility heuristic, and the three rejected alternatives.

Carry forwardCost automation on reclaimable capacity needs an availability term, even a crude one, before it is allowed to act.
github.com/kubernetes-sigs/karpenter/designs/spot-consolidation.md
Design record Kubernetes SIG Node2020-10

KEP-2000: Graceful node shutdown

The kubelet grew a shutdown lifecycle because preemptible VMs kept killing pods without one. Systemd inhibitor locks delay the shutdown; a grace budget is split between ordinary and critical pods. Defaults to zero seconds until configured.

Carry forwardThe vacate path exists but is off by default outside managed offerings; verify yours is actually configured.
github.com/kubernetes/enhancements/keps/sig-node/2000
Design record Karpenter2025

Capacity buffers RFC

Formalises paid headroom (a CapacityBuffer API with virtual pods in the scheduling simulation) after years of balloon-pod and static-NodePool workarounds, requested across at least five community issues.

Carry forwardJust-in-time capacity converges back toward paying for slack; budget the buffer as part of the spot business case.
github.com/kubernetes-sigs/karpenter/designs/capacity-buffers.md
Source Karpenter user2024-03

Issue #1124: "requires 15 cheaper instance type options, got 1"

The availability guard as users experience it: consolidation refuses to run, the documentation said "15 instance types" while the code wants 15 cheaper ones, and the discrepancy reads as a bug.

Carry forwardGuardrails that silently withhold savings get reported as defects; surface the guard's reasoning in the tool's own output.
github.com/kubernetes-sigs/karpenter/issues/1124
Source Karpenter user2024-09

Issue #1653: a t3.2xlarge at 26% allocation that will not consolidate

Six cheaper types available, nine short of the floor; the node idles at 26% CPU and 7% memory allocation. Maintainers labelled it kind/feature and priority/awaiting-more-evidence rather than relaxing the guard.

Carry forwardThe cost of the availability guard is real and visible per node; price it against the interruption churn it prevents.
github.com/kubernetes-sigs/karpenter/issues/1653
Source Kubernetes autoscaler user2020-09

Issue #3490: priority expander vs unavailable spot

The fallback path under real scarcity: spot group unhealthy only after ~10 minutes, placeholder instances wedging group selection, on-demand group never engaged within the configured window. Closed rotten, unresolved.

Carry forwardSerial group-by-group retry is structurally slow; prefer launch requests that carry the whole acceptable set at once.
github.com/kubernetes/autoscaler/issues/3490
Source AWSchecked 2026-09

aws-node-termination-handler

The two signal architectures (per-node IMDS polling; central SQS/EventBridge queue with lifecycle hooks and heartbeats up to 48 hours), and the boundary note that EKS managed node groups make the tool unnecessary.

Carry forwardFleet-scale spot needs the queue architecture; per-node polling is for small estates.
github.com/aws/aws-node-termination-handler
Source NTH user2021-02

Issue #368: drain on rebalance recommendation

The argument for acting early, in the user's own words: two minutes is not enough when the cluster must first create a node to receive the drained pods. The feature shipped as opt-in; Karpenter later declined the same behaviour.

Carry forwardAct on the early signal only if your measured drain-plus-replace exceeds the notice window.
github.com/aws/aws-node-termination-handler/issues/368
Vendor AWS2023 archive

EC2 Spot best practices (documentation source)

Up to 90%; at least 10 instance types; price-capacity-optimized recommended; and the strong warning against spot, and against on-demand failover, for workloads intolerant of incomplete target capacity.

Carry forwardThe vendor's own docs refuse the availability story most spot adopters tell themselves; read the warning before quoting the 90%.
github.com/awsdocs/amazon-ec2-user-guide/…/spot-best-practices.md
Vendor AWS2023 archive

Spot interruption notices (documentation source)

The contract's fine print: two minutes, poll every five seconds, delivered as CloudWatch event and instance metadata, and "emitted on a best effort basis". Hibernation forfeits the lead time.

Carry forwardDesign the vacate path to survive receiving no notice at all; the warning is a courtesy, not a guarantee.
github.com/awsdocs/amazon-ec2-user-guide/…/spot-instance-termination-notices.md
Vendor AWS2023 archive

Rebalance recommendations (documentation source)

The early-warning signal's honest specification: it is not always possible to send it before the interruption notice, so it "can arrive along with the two-minute interruption notice". Only instances launched after November 2020 receive it.

Carry forwardTreat the rebalance recommendation as a hint for scheduling decisions, never as the trigger for guaranteed-safe migration.
github.com/awsdocs/amazon-ec2-user-guide/…/rebalance-recommendations.md
Vendor AWS2023 archive

Spot Instances overview and reasons for interruption

Post-2017 pricing: set by EC2, "adjusted gradually based on the long-term supply of and demand". Reclaim reasons: capacity first ("when it needs it back"), price only if you cap it, and capping raises interruption frequency.

Carry forwardModel interruptions as a capacity phenomenon, not a price phenomenon; the bidding-war era ended in 2017.
github.com/awsdocs/amazon-ec2-user-guide/…/using-spot-instances.md
Vendor Microsoft2026-02

About Azure Spot Virtual Machines

30-second best-effort scheduled events, "no SLA", "no high availability guarantees"; eviction rates quoted per hour from seven days of history; deallocate-vs-delete policies, with reallocation after deallocate explicitly not guaranteed.

Carry forwardAzure's per-hour eviction bands are the only provider-published interruption-rate numbers reachable in this corpus; use their unit when you define your own SLOs.
github.com/MicrosoftDocs/azure-compute-docs/…/spot-vms.md
Vendor Google Cloud2022–2026

Spot VMs product page and GA announcement

Up to 91% off (60–91% at GA), 30 seconds to shut down, GKE wiring preemption into kubelet graceful shutdown, and the caveat that carries this guide's thesis: replacement happens "if capacity is available".

Carry forwardAutomatic recreation is conditional; the condition is the part to engineer for.
cloud.google.com/solutions/spot-vms
Vendor AWS / Karpenterchecked 2026-09

Karpenter disruption documentation

Interruption handling built in: taint, drain and replace on the two-minute notice, with replacement launched in parallel; rebalance recommendations surfaced as events but deliberately not acted on.

Carry forwardDrain and replace concurrently; the notice window is too short for sequential handling.
github.com/aws/karpenter-provider-aws/…/disruption.md
Eng blog Datadog2026-03

Understanding Karpenter architecture

The capacity-type priority order (reserved, spot, on-demand) and the admission rule: interruption-intolerant workloads should be explicitly pinned off spot by NodePool requirements.

Carry forwardMake riding spot an explicit per-workload opt-in encoded in scheduling constraints.
datadoghq.com/blog/karpenter-architecture
Eng blog Datadog2026-03

Key metrics for monitoring Karpenter

The churn dashboard: interruption message counts, disruption reasons, termination duration, and pairing interruption rate with pod startup duration to test whether interruptions remain "transparent cost optimization".

Carry forwardAlert on the churn budget and its workload-visible effects, not on individual reclaim events.
datadoghq.com/blog/karpenter-key-metrics
Paper SkyPilot / UC Berkeley2024

spot-traces: artifact of "Can't Be Late" (NSDI '24, Outstanding Paper)

Multi-week spot availability and preemption traces on AWS and GCP, collected by directly pinging the providers; up to ~45 continuous days for V100 pools, single- and multi-node. The USENIX paper itself was unreachable from this session; the artifact and its data are on GitHub.

Carry forwardPreemption behaviour is measurable from the outside; before betting a workload on a pool, trace it the same way.
github.com/skypilot-org/spot-traces
07

Build a miniature, then productionise it

Six rungs from a single interruptible VM to a rehearsed capacity incident. The line from toy to real is crossed at rung 4, where you stop trusting the fallback and start measuring it.

Catch one repossession with your own eyes

Launch one spot/Spot VM, run a loop polling the provider's interruption endpoint (every 5 seconds on AWS, per the docs), and leave it until it dies or you trigger the provider's interruption-simulation tooling. Log the notice-to-termination gap.

Done when: you have a timestamped log showing notice receipt and actual termination.  Teaches: the notice is short, best-effort, and arrives on a plain metadata endpoint, not somewhere exotic.

Make a workload vacate inside the window

Run a service with graceful shutdown on that node; on notice, cordon, drain and verify in-flight work completes. Size terminationGracePeriodSeconds to the 30-second providers, not the 120-second one.

Done when: zero failed requests across ten simulated interruptions.  Teaches: KEP-2000-style shutdown handling defaults to off; the grace budget is yours to configure and verify.

Diversify until the guard would let you consolidate

Express the workload's requirements so at least 10–15 instance types across every zone satisfy them (AWS's rule of thumb and Karpenter's enforced floor). Launch through a capacity-aware allocation strategy rather than a type list you hand-rank.

Done when: your launch request names ≥10 acceptable types and you have seen the allocator pick a non-cheapest one.  Teaches: flexibility is the availability control; the allocator's job is to spend it wisely.

Starve it: rehearse the empty pool

Constrain the fleet to a deliberately narrow pool set, generate load, and take that pool away (quota, exclusion, or an off-peak zone drain). Measure signal-to-serving time for the fallback, and watch for the #3490 pattern: the primary group retried serially while pods sit pending.

Done when: you can state your fallback's engage time in minutes from measurement, not from configuration.  Teaches: a configured timeout is not an observed one; GitLab's incident review shows even the firm tier needs this rehearsal.

Meter the churn and give it a budget

Export interruption counts, node lifetime, and pod startup duration; define a weekly churn budget per workload and alert on the budget. Trace one pool for a week, spot-traces style, and compare its behaviour with the provider's published band.

Done when: you can answer "what did interruptions cost us last week, in replacement latency" from a dashboard.  Teaches: the discount's true price is churn, and churn is measurable.

Run the capacity game day

Peak load, primary pools removed, on-demand quota capped to a realistic ceiling. The team must choose among queueing, shedding and paying, using the runbook, then adopt one stateless production workload onto spot GitLab-style: reviewed change, named rollback metric, revert path.

Done when: the game day produces a written decision record and the production adoption survives its first real interruption without a page.  Teaches: the difference between tolerating an interruption and tolerating a shortage, which is the distinction AWS's warning turns on.

08

Keep hunting

The queries that found this material. The provider-error vocabulary (ZONE_RESOURCE_POOL_EXHAUSTED, InsufficientInstanceCapacity, SpotToSpotConsolidation) is the highest-yield thread: error strings appear verbatim in incident reviews and issue trackers.

Production experience and incidents

  • "spot instances" "lessons learned" production -tutorial "interruption"
  • "spot interruption" postmortem OR "incident review" "we"
  • ZONE_RESOURCE_POOL_EXHAUSTED site:gitlab.com
  • "insufficient capacity" OR stockout gl-infra production incident

The design arguments

  • repo:kubernetes-sigs/karpenter path:designs spot
  • SpotToSpotConsolidation issue "15 cheaper"
  • "rebalance recommendation" drain churn issue
  • cluster-autoscaler priority expander spot fallback "max-node-provision-time"

The contracts themselves

  • site:github.com awsdocs spot-best-practices
  • site:github.com MicrosoftDocs spot-vms eviction
  • spot VMs "if capacity is available" site:cloud.google.com

Data and papers

  • spot availability preemption traces github dataset
  • "can't be late" spot deadlines nsdi
  • spot instance eviction rate "per hour" measurement
09

References

  1. Karpenter, Spot Consolidation design doc kubernetes-sigs/karpenter, merged for v0.34 (2024). Checked 2026-09-29.
  2. Karpenter issue #1124, SpotToSpotConsolidation requires 15 cheaper instance type options, got 1 GitHub, 2024-03-21. Checked 2026-09-29.
  3. Karpenter issue #1653, SpotToSpotConsolidation … got 6 GitHub, 2024-09-10. Checked 2026-09-29.
  4. Karpenter, Capacity Buffer Support RFC kubernetes-sigs/karpenter. Checked 2026-09-29.
  5. kubernetes/autoscaler issue #3490, Max Node Provision Time + Priority Expander + Node Unavailability GitHub, 2020-09-04. Checked 2026-09-29.
  6. AWS, aws-node-termination-handler README GitHub. Checked 2026-09-29.
  7. aws-node-termination-handler issue #368, Drain on Rebalance Recommendation Notification GitHub, 2021-02-18. Checked 2026-09-29.
  8. Kubernetes SIG Node, KEP-2000: Graceful Node Shutdown kubernetes/enhancements, approved 2020-10-02. Checked 2026-09-29.
  9. Karpenter documentation, Disruption (interruption handling) aws/karpenter-provider-aws. Checked 2026-09-29.
  10. AWS, Best practices for EC2 Spot (documentation source) awsdocs/amazon-ec2-user-guide, repository archived June 2023. Checked 2026-09-29.
  11. AWS, Spot Instance interruption notices (documentation source) awsdocs/amazon-ec2-user-guide, archived June 2023. Checked 2026-09-29.
  12. AWS, EC2 instance rebalance recommendations (documentation source) awsdocs/amazon-ec2-user-guide, archived June 2023. Checked 2026-09-29.
  13. AWS, Reasons for interruption (documentation source) awsdocs/amazon-ec2-user-guide, archived June 2023. Checked 2026-09-29.
  14. AWS, Spot Instances overview (documentation source) awsdocs/amazon-ec2-user-guide, archived June 2023. Checked 2026-09-29.
  15. Microsoft, About Azure Spot Virtual Machines (documentation source) MicrosoftDocs/azure-compute-docs, ms.date 2026-02-06. Checked 2026-09-29.
  16. Google Cloud, Spot VMs product page cloud.google.com. Checked 2026-09-29.
  17. Google Cloud, Spot VMs now GA Google Cloud blog, 2022-05-20. Checked 2026-09-29.
  18. GitLab, Post-Incident Review INC-9343 (ci_runner_jobs SLO violation, saas-linux-small-amd64) GitLab production tracker, 2026-04-20. Checked 2026-09-29.
  19. GitLab, INC-9866: Frequent GKE node scale-up errors due to machine type shortage GitLab production tracker, 2026-05-05. Checked 2026-09-29.
  20. GitLab, GKE unable to scale due to lack of SSD availability GitLab production tracker, 2021-06-21. Checked 2026-09-29.
  21. GitLab, GKE node pool constraints: N4 vCPU limits + N4D stockouts GitLab production tracker, 2026-08-07. Checked 2026-09-29.
  22. GitLab, Change: allow Mimir distributors on spot instances GitLab production tracker, 2024-10-22. Checked 2026-09-29.
  23. Datadog, Understanding Karpenter architecture for Kubernetes autoscaling Datadog blog, 2026-03-11. Checked 2026-09-29.
  24. Datadog, Key metrics for monitoring Karpenter Datadog blog, 2026-03-11. Checked 2026-09-29.
  25. SkyPilot, spot-traces (artifact of "Can't Be Late", NSDI '24) GitHub, 2024. Checked 2026-09-29.

Known but unreachable from this session's network, therefore uncited: AWS's November 2017 announcement replacing spot bidding with smoothed pricing; Netflix's "Creating Your Own EC2 Spot Market" (2015); Delivery Hero's and Honeycomb's spot posts; the full text of "Can't Be Late" on usenix.org; AWS Spot Instance Advisor. The evidence ledger (sources.md) records each as a named gap.