Every source behind this page, graded. The mix is honest about its
constraint: this session's network could reach GitHub, GitLab, Google Cloud and Datadog,
so the classic engineering-blog accounts (Netflix's internal spot market, Delivery Hero,
Honeycomb) are named in the ledger as unfetched rather than cited here.
Postmortem
GitLab2026-04
Incident review INC-9343: CI runner shard SLO violation
A GCP zonal capacity failure surfaced as ZONE_RESOURCE_POOL_EXHAUSTED; all six runner
managers were pinned to the failing zone. 2h55m, apdex to ~20%, 24K queued jobs.
Operators diagnosed the zone ~3 hours before the provider's public acknowledgement.
Carry forwardZone diversification that the launch path does
not exercise is decorative. Keep a canary launch as ground truth.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/21838
Postmortem
GitLab2026-05
INC-9866: GKE node scale-up errors from machine-type shortage
Scale-ups failed at up to 77.2% over 24 hours because the provider had misallocated
newly delivered machines between product pools. The customer's fix was family-level
fallback node pools.
Carry forwardThe firm tier stocks out too; machine-family
flexibility is an availability control, not just a spot tactic.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/21996
Postmortem
GitLab2021-06
GKE unable to scale: SSD availability
An earlier instance of the same class: 146 minutes of impact when local-SSD-backed
nodes could not be created in-zone. Same failure shape five years before INC-9343.
Carry forwardStockouts are a recurring class, not a freak
event; design and drill for them on that basis.
gitlab.com/gitlab-com/gl-infra/production/-/issues/4940
Case study
GitLab2026-08
N4 quota limits and N4D stockouts blocking deploys
Three consecutive production deploys slowed to 45 minutes by zonal stockouts; the
investigation explicitly feeds machine-family strategy into a GCP contract negotiation
before per-SKU discounts freeze.
Carry forwardDecide instance-family flexibility before the
discount contract is signed, not after.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/22668
Case study
GitLab2024-10
Change record: Mimir distributors allowed onto spot
How a mature operator actually adopts spot: one stateless service, a reviewed change,
a named rollback metric (Mimir write latency) and a five-minute revert path.
Carry forwardAdopt per workload, with the rollback metric
written down before the change ships.
gitlab.com/gitlab-com/gl-infra/production/-/work_items/18749
Design record
Karpenter2023/24
Spot consolidation design
The central ADR of this topic: why pure price minimization on spot self-destructs
("walking down the PCO decision ladder"), why availability must be proxied by a
flexibility heuristic, and the three rejected alternatives.
Carry forwardCost automation on reclaimable capacity needs
an availability term, even a crude one, before it is allowed to act.
github.com/kubernetes-sigs/karpenter/designs/spot-consolidation.md
Design record
Kubernetes SIG Node2020-10
KEP-2000: Graceful node shutdown
The kubelet grew a shutdown lifecycle because preemptible VMs kept killing pods
without one. Systemd inhibitor locks delay the shutdown; a grace budget is split
between ordinary and critical pods. Defaults to zero seconds until configured.
Carry forwardThe vacate path exists but is off by default
outside managed offerings; verify yours is actually configured.
github.com/kubernetes/enhancements/keps/sig-node/2000
Design record
Karpenter2025
Capacity buffers RFC
Formalises paid headroom (a CapacityBuffer API with virtual pods in the scheduling
simulation) after years of balloon-pod and static-NodePool workarounds, requested
across at least five community issues.
Carry forwardJust-in-time capacity converges back toward
paying for slack; budget the buffer as part of the spot business case.
github.com/kubernetes-sigs/karpenter/designs/capacity-buffers.md
Source
Karpenter user2024-03
Issue #1124: "requires 15 cheaper instance type options, got 1"
The availability guard as users experience it: consolidation refuses to run, the
documentation said "15 instance types" while the code wants 15 cheaper ones,
and the discrepancy reads as a bug.
Carry forwardGuardrails that silently withhold savings get
reported as defects; surface the guard's reasoning in the tool's own output.
github.com/kubernetes-sigs/karpenter/issues/1124
Source
Karpenter user2024-09
Issue #1653: a t3.2xlarge at 26% allocation that will not consolidate
Six cheaper types available, nine short of the floor; the node idles at 26% CPU and
7% memory allocation. Maintainers labelled it kind/feature and
priority/awaiting-more-evidence rather than relaxing the guard.
Carry forwardThe cost of the availability guard is real and
visible per node; price it against the interruption churn it prevents.
github.com/kubernetes-sigs/karpenter/issues/1653
Source
Kubernetes autoscaler user2020-09
Issue #3490: priority expander vs unavailable spot
The fallback path under real scarcity: spot group unhealthy only after ~10 minutes,
placeholder instances wedging group selection, on-demand group never engaged within
the configured window. Closed rotten, unresolved.
Carry forwardSerial group-by-group retry is structurally
slow; prefer launch requests that carry the whole acceptable set at once.
github.com/kubernetes/autoscaler/issues/3490
Source
AWSchecked 2026-09
aws-node-termination-handler
The two signal architectures (per-node IMDS polling; central SQS/EventBridge queue
with lifecycle hooks and heartbeats up to 48 hours), and the boundary note that EKS
managed node groups make the tool unnecessary.
Carry forwardFleet-scale spot needs the queue architecture;
per-node polling is for small estates.
github.com/aws/aws-node-termination-handler
Source
NTH user2021-02
Issue #368: drain on rebalance recommendation
The argument for acting early, in the user's own words: two minutes is not enough
when the cluster must first create a node to receive the drained pods. The feature
shipped as opt-in; Karpenter later declined the same behaviour.
Carry forwardAct on the early signal only if your measured
drain-plus-replace exceeds the notice window.
github.com/aws/aws-node-termination-handler/issues/368
Vendor
AWS2023 archive
EC2 Spot best practices (documentation source)
Up to 90%; at least 10 instance types; price-capacity-optimized recommended; and the
strong warning against spot, and against on-demand failover, for workloads intolerant
of incomplete target capacity.
Carry forwardThe vendor's own docs refuse the availability
story most spot adopters tell themselves; read the warning before quoting the 90%.
github.com/awsdocs/amazon-ec2-user-guide/…/spot-best-practices.md
Vendor
AWS2023 archive
Spot interruption notices (documentation source)
The contract's fine print: two minutes, poll every five seconds, delivered as
CloudWatch event and instance metadata, and "emitted on a best effort basis".
Hibernation forfeits the lead time.
Carry forwardDesign the vacate path to survive receiving no
notice at all; the warning is a courtesy, not a guarantee.
github.com/awsdocs/amazon-ec2-user-guide/…/spot-instance-termination-notices.md
Vendor
AWS2023 archive
Rebalance recommendations (documentation source)
The early-warning signal's honest specification: it is not always possible to send it
before the interruption notice, so it "can arrive along with the two-minute
interruption notice". Only instances launched after November 2020 receive it.
Carry forwardTreat the rebalance recommendation as a hint
for scheduling decisions, never as the trigger for guaranteed-safe migration.
github.com/awsdocs/amazon-ec2-user-guide/…/rebalance-recommendations.md
Vendor
AWS2023 archive
Spot Instances overview and reasons for interruption
Post-2017 pricing: set by EC2, "adjusted gradually based on the long-term supply of
and demand". Reclaim reasons: capacity first ("when it needs it back"), price only if
you cap it, and capping raises interruption frequency.
Carry forwardModel interruptions as a capacity phenomenon,
not a price phenomenon; the bidding-war era ended in 2017.
github.com/awsdocs/amazon-ec2-user-guide/…/using-spot-instances.md
Vendor
Microsoft2026-02
About Azure Spot Virtual Machines
30-second best-effort scheduled events, "no SLA", "no high availability guarantees";
eviction rates quoted per hour from seven days of history; deallocate-vs-delete
policies, with reallocation after deallocate explicitly not guaranteed.
Carry forwardAzure's per-hour eviction bands are the only
provider-published interruption-rate numbers reachable in this corpus; use their unit
when you define your own SLOs.
github.com/MicrosoftDocs/azure-compute-docs/…/spot-vms.md
Vendor
Google Cloud2022–2026
Spot VMs product page and GA announcement
Up to 91% off (60–91% at GA), 30 seconds to shut down, GKE wiring preemption
into kubelet graceful shutdown, and the caveat that carries this guide's thesis:
replacement happens "if capacity is available".
Carry forwardAutomatic recreation is conditional; the
condition is the part to engineer for.
cloud.google.com/solutions/spot-vms
Vendor
AWS / Karpenterchecked 2026-09
Karpenter disruption documentation
Interruption handling built in: taint, drain and replace on the two-minute notice,
with replacement launched in parallel; rebalance recommendations surfaced as events
but deliberately not acted on.
Carry forwardDrain and replace concurrently; the notice
window is too short for sequential handling.
github.com/aws/karpenter-provider-aws/…/disruption.md
Eng blog
Datadog2026-03
Understanding Karpenter architecture
The capacity-type priority order (reserved, spot, on-demand) and the admission rule:
interruption-intolerant workloads should be explicitly pinned off spot by NodePool
requirements.
Carry forwardMake riding spot an explicit per-workload
opt-in encoded in scheduling constraints.
datadoghq.com/blog/karpenter-architecture
Eng blog
Datadog2026-03
Key metrics for monitoring Karpenter
The churn dashboard: interruption message counts, disruption reasons, termination
duration, and pairing interruption rate with pod startup duration to test whether
interruptions remain "transparent cost optimization".
Carry forwardAlert on the churn budget and its
workload-visible effects, not on individual reclaim events.
datadoghq.com/blog/karpenter-key-metrics
Paper
SkyPilot / UC Berkeley2024
spot-traces: artifact of "Can't Be Late" (NSDI '24, Outstanding Paper)
Multi-week spot availability and preemption traces on AWS and GCP, collected by
directly pinging the providers; up to ~45 continuous days for V100 pools, single- and
multi-node. The USENIX paper itself was unreachable from this session; the artifact
and its data are on GitHub.
Carry forwardPreemption behaviour is measurable from the
outside; before betting a workload on a pool, trace it the same way.
github.com/skypilot-org/spot-traces