Evidence ledger
One row per claim in The discount is a repossession clause: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Topic: how production systems run on interruptible (spot / preemptible) cloud capacity — the interruption contract, the machinery that absorbs repossession, and where the discount stops paying for it.
Network note: this session's egress allowlist permitted github.com, raw.githubusercontent.com, gitlab.com, cloud.google.com, www.datadoghq.com and package registries. Engineering blogs (Netflix, Delivery Hero, Honeycomb), AWS's own site, USENIX, arXiv and conference video hosts were unreachable and are therefore absent or represented through GitHub-hosted artifacts. That constraint is stated in the guide.
| # | Org | Title | Tier | Published | Checked | URL | Claim I take from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | Karpenter (kubernetes-sigs) | Spot Consolidation design doc | adr | 2023 (merged for v0.34, 2024-01) | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/blob/main/designs/spot-consolidation.md | Pure price minimization on spot self-destructs; consolidation was withheld until an availability heuristic existed | "If Karpenter were to consolidate spot capacity based purely on cost, it would repeatedly interrupt the same instance, walking down the PCO decision ladder, until only the lowest price remained." |
| 2 | Karpenter (kubernetes-sigs) | Spot Consolidation design doc | adr | 2023/24 | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/blob/main/designs/spot-consolidation.md | Availability is invisible to the customer, so the guardrail must be a heuristic | "cloud providers don't surface capacity pool information (availability) for spot, meaning that we need to define availability by some heuristic to make this optimization work" |
| 3 | Karpenter (kubernetes-sigs) | Spot Consolidation design doc | adr | 2023/24 | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/blob/main/designs/spot-consolidation.md | The 15-cheaper-types floor and its origin; rejected options 2-4 (price improvement factor, optimistic launches, minimum node lifetime) | "Minimum flexibility of 15 spot instance types is decided after an analysis done on the flexibility of AWS customers request in the launch path." |
| 4 | Karpenter user (1ms-ms) | Issue #1124: SpotToSpotConsolidation requires 15 cheaper instance types … got 1 | source | 2024-03-21 | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/issues/1124 | Users hit the guardrail and read it as a bug; docs said "15 instance types", code wants 15 cheaper ones | "Why does it require 15 cheaper nodes? As per docs it requires 15 instance types, not 15 cheaper instance types." |
| 5 | Karpenter user (dmitry-mightydevops) | Issue #1653: SpotToSpotConsolidation … got 6 | source | 2024-09-10 | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/issues/1653 | A t3.2xlarge at 26% CPU allocation is left running because only 6 cheaper types fit the NodePool; labelled kind/feature, priority/awaiting-more-evidence | Error: "SpotToSpotConsolidation requires 15 cheaper instance type options than the current candidate to consolidate, got 6"; node usage "~26% CPU (2080m) and 7% memory" |
| 6 | Kubernetes autoscaler user (nitrag) | Issue #3490: Max Node Provision Time + Priority Expander + Node Unavailability | source | 2020-09-04 | 2026-09-29 | https://github.com/kubernetes/autoscaler/issues/3490 | The spot→on-demand fallback path has minutes of latency and can wedge on placeholder instances; closed rotten, unresolved | "Instance group pv-platform-prod-worker-gpu-accelerated-p2-xlarge-spot has only 0 instances created while requested count is 1. Creating placeholder instance…"; unhealthy at ~10 min despite a 5-min max-node-provision-time |
| 7 | AWS | aws-node-termination-handler README | source | current (checked) | 2026-09-29 | https://github.com/aws/aws-node-termination-handler | The two architectures for hearing the interruption signal: IMDS DaemonSet polling vs EventBridge→SQS queue processor; handles spot ITN, rebalance recommendations, ASG lifecycle | README documents both modes; queue processor supports lifecycle heartbeats "up to 48 hours"; "If you're using EKS managed node groups, you don't need the aws-node-termination-handler." |
| 8 | AWS NTH user (dongheeJeong) | Issue #368: Drain on Rebalance Recommendation Notification | source | 2021-02-18 | 2026-09-29 | https://github.com/aws/aws-node-termination-handler/issues/368 | Why users wanted to act on the early-warning signal: two minutes is not enough when the cluster must first make room | "Two minutes may not sufficient time to drain node on some circumstance. (e.g. When new worker node creation needed when all existing worker nodes are too busy to accept pods)." |
| 9 | Kubernetes SIG Node | KEP-2000: Graceful Node Shutdown | adr | approved 2020-10-02, alpha v1.20 | 2026-09-29 | https://github.com/kubernetes/enhancements/blob/master/keps/sig-node/2000-graceful-node-shutdown/README.md | Spot/preemptible reshaped the kubelet itself; the mechanism is a systemd inhibitor lock; defaults to off (0s) | Motivation names shutdown "controlled by some external system (e.g. Preemptible VMs)"; kubelet takes a "delay systemd inhibitor lock"; ShutdownGracePeriod "Defaults to 0 seconds; requires explicit configuration to activate" |
| 10 | Karpenter (aws-provider) | Capacity Buffers RFC | adr | 2025 (in main, checked) | 2026-09-29 | https://github.com/kubernetes-sigs/karpenter/blob/main/designs/capacity-buffers.md | The ecosystem is re-adding paid headroom on top of just-in-time provisioning; balloon-pod workarounds get formalized into a CapacityBuffer API | "Users maintain spare capacity through balloon pods, pause containers, or NodePools with fixed capacity"; feature "requested multiple times by the community" (5 linked issues) |
| 11 | AWS | EC2 User Guide: Spot best practices (docs source, archived repo) | vendor | repo archived 2023-06 | 2026-09-29 | https://github.com/awsdocs/amazon-ec2-user-guide/blob/master/doc_source/spot-best-practices.md | Up to 90% claim; ≥10 instance types rule of thumb; AWS itself warns against spot + on-demand failover for capacity-intolerant workloads | "savings of up to 90% off"; "A good rule of thumb is to be flexible across at least 10 instance types"; "not recommended for workloads that are intolerant of occasional periods when the target capacity is not completely available. We strongly warn against using Spot Instances for these workloads or attempting to fail-over to On-Demand Instances to handle interruptions." |
| 12 | AWS | EC2 User Guide: Spot Instance interruption notices (docs source) | vendor | repo archived 2023-06 | 2026-09-29 | https://github.com/awsdocs/amazon-ec2-user-guide/blob/master/doc_source/spot-instance-termination-notices.md | The two-minute warning is itself best-effort; poll every 5 s; hibernation gets no two-minute lead | "a warning that is issued two minutes before"; "We recommend that you check for these interruption notices every 5 seconds"; "Interruption notices are emitted on a best effort basis." |
| 13 | AWS | EC2 User Guide: rebalance recommendations (docs source) | vendor | repo archived 2023-06 | 2026-09-29 | https://github.com/awsdocs/amazon-ec2-user-guide/blob/master/doc_source/rebalance-recommendations.md | The early warning may not be early | "It is not always possible for Amazon EC2 to send the rebalance recommendation signal before the two-minute Spot Instance interruption notice. Therefore, the rebalance recommendation signal can arrive along with the two-minute interruption notice." |
| 14 | AWS | EC2 User Guide: reasons for interruption (docs source) | vendor | repo archived 2023-06 | 2026-09-29 | https://github.com/awsdocs/amazon-ec2-user-guide/blob/master/doc_source/interruption-reasons.md | Capacity, not price, is the operative reclaim reason post-2017; setting a max price raises interruption frequency | "Amazon EC2 can interrupt your Spot Instance when it needs it back. EC2 reclaims your instance mainly to repurpose capacity"; "if you specify a maximum price, your instances will be interrupted more frequently than if you do not specify it" |
| 15 | Microsoft | About Azure Spot Virtual Machines (docs source) | vendor | ms.date 2026-02-06 | 2026-09-29 | https://github.com/MicrosoftDocs/azure-compute-docs/blob/main/articles/virtual-machines/spot-vms.md | 30-second best-effort notice, no SLA, eviction rates quoted per HOUR from 7-day history; deallocate-vs-delete policy trade | "no SLA … no high availability guarantees … evict Azure Spot Virtual Machines with 30-seconds notice"; scheduled events "delivered on a best effort basis up to 30 seconds prior"; "an eviction rate of 10% means a VM has a 10% chance of being evicted within the next hour, based on historical eviction data of the last 7 days"; deallocated VMs: "no guarantee that the allocation will succeed" |
| 16 | Google Cloud | Spot VMs product page | vendor | current (checked) | 2026-09-29 | https://cloud.google.com/solutions/spot-vms | Up to 91% off; 30 seconds to shut down; replacement is conditional on capacity | "reduce your Compute Engine costs by up to 91%"; "Compute Engine gives you 30 seconds to shut down when you're preempted"; "Managed instance groups automatically recreate your instances when they're preempted (if capacity is available)." |
| 17 | Google Cloud | Spot VMs GA announcement | vendor | 2022-05-20 | 2026-09-29 | https://cloud.google.com/blog/products/compute/google-cloud-spot-vms-now-ga | 60–91% discount band at GA; GKE wires preemption into kubelet graceful node shutdown | "60-91% discount"; "the kubelet graceful node shutdown feature is enabled by default, which allows the kubelet to notice the preemption notice, gracefully terminate Pods…" |
| 18 | AWS (Karpenter docs) | Karpenter disruption documentation | vendor | current (checked) | 2026-09-29 | https://github.com/aws/karpenter-provider-aws/blob/main/website/content/en/docs/concepts/disruption.md | Karpenter drains and pre-provisions on the 2-minute ITN but deliberately does NOT act on rebalance recommendations — opposite call to NTH | "Karpenter's average node startup time means that, generally, there is sufficient time for the new node to become ready"; "Karpenter does not currently support taint, drain, and terminate logic for Spot Rebalance Recommendations." |
| 19 | GitLab | Post-Incident Review INC-9343: ci_runner_jobs SLO violation on saas-linux-small-amd64 | postmortem | 2026-04-20 | 2026-09-29 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/21838 | A zonal capacity event (ZONE_RESOURCE_POOL_EXHAUSTED on n2d-standard-2) exhausted a runner shard: apdex ~100%→~20%, 24K jobs queued, 2h55m; all six runner managers were pinned to one zone | "Because all six d-pinned runner managers on saas-linux-small-amd64 could not create new VMs, CI runner capacity was exhausted"; "apdex dropped from ~100% to ~20% at the worst point…; pending queue peaked around 24K jobs"; duration "2 hours, 55 minutes"; GitLab diagnosed zone-local failure ~3h before GCP's public acknowledgement |
| 20 | GitLab | INC-9866: Frequent GKE node scale-up errors due to machine type shortage | postmortem | 2026-05-05 | 2026-09-29 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/21996 | On-demand machine types stock out too; scale-up failure up to 77.2% over 24h; provider misallocated new machines between product pools | "failed at rates up to 77.2% over the past 24 hours"; "Google confirmed that newly delivered machines were misallocated into a different product pool, leading to resource stockouts." |
| 21 | GitLab | GKE unable to scale due to lack of SSD availability | postmortem | 2021-06-21 | 2026-09-29 | https://gitlab.com/gitlab-com/gl-infra/production/-/issues/4940 | Zonal resource stockouts are a recurring, years-long class, not a one-off | Incident record: customer impact duration "08:10 - 10:36 (146 minutes)", root cause label Service::GCP |
| 22 | GitLab | Investigate GKE node pool constraints (N4 vCPU limits + N4D stockouts) | casestudy | 2026-08-07 | 2026-09-29 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/22668 | Stockouts turned machine-family choice into a contract-negotiation input; deploys waited 45 min; the binding constraint was provider capacity, not workload size | "a single deploy job recently taking 45 minutes"; "Scale-ups on N4 fail against a vCPU quota limit, and scale-ups on N4D fail against zonal stockouts"; "the binding constraint is node availability, not the size of our own workload" |
| 23 | GitLab | Change record: allow Mimir distributors on spot instances | casestudy | 2024-10-22 | 2026-09-29 | https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/18749 | Spot adoption in a mature org is per-workload, reviewed, reversible, with named rollback metrics | "updates the Mimir distributors to allow scheduling on spot instances… prefer scheduling on spot instances during scaling operations during peak hours"; rollback trigger: "undesirable increase in Mimir write latency" |
| 24 | Datadog | Understanding Karpenter architecture | blog | 2026-03-11 | 2026-09-29 | https://www.datadoghq.com/blog/karpenter-architecture/ | Priority ordering reserved→spot→on-demand; workload admission is the operator's job | "For workloads that cannot tolerate the potential interruptions associated with Spot capacity, cluster administrators should explicitly limit their NodePool requirements to on-demand or reserved types" |
| 25 | Datadog | Key metrics for monitoring Karpenter | blog | 2026-03-11 | 2026-09-29 | https://www.datadoghq.com/blog/karpenter-key-metrics/ | Churn is the observable of interest: interruption counter plus pod-startup-duration tells you whether interruptions are "transparent cost optimization" or damage | Metric karpenter_interruption_received_messages_total; "This increased rate of node replacements can create brief capacity gaps"; "A corresponding rise in karpenter_pods_startup_duration_seconds confirms that these interruptions aren't just transparent cost optimization" |
| 26 | SkyPilot (UC Berkeley) | spot-traces dataset (artifact of "Can't Be Late", NSDI '24 Outstanding Paper) | paper | 2024 | 2026-09-29 | https://github.com/skypilot-org/spot-traces | Preemption behaviour is measurable and was measured: multi-week availability/preemption traces on AWS and GCP, collected "by directly pinging the cloud providers"; longest continuous trace ~45 days (V100) | README: traces spanning 2 days to 2 months across AWS/GCP GPU and CPU types, us-east-1/us-east-2/us-west-2/us-central1, single- and multi-node up to 16 nodes |
Unreachable but known material (named gaps, no citations used)
- AWS's 2017 removal of spot bidding ("New Amazon EC2 Spot pricing model", Nov 2017) — aws.amazon.com unreachable from this session; the post-2017 mechanism is cited instead from the EC2 User Guide docs source (rows 12, 14: prices "adjusted gradually based on the long-term supply of and demand").
- Netflix "Creating Your Own EC2 Spot Market" (2015), Delivery Hero spot posts, Honeycomb instance-type posts — engineering-blog hosts blocked; found via search, not fetched, not cited.
- "Can't Be Late" full paper text (USENIX) — only its GitHub artifact was reachable (row 26).
- AWS Spot Instance Advisor interruption-frequency bands — aws.amazon.com unreachable; Azure's per-hour eviction-rate definition (row 15) is the citable equivalent.
| 27 | AWS | EC2 User Guide: Spot Instances overview (docs source) | vendor | repo archived 2023-06 | 2026-09-29 | https://github.com/awsdocs/amazon-ec2-user-guide/blob/master/doc_source/using-spot-instances.md | Post-2017 prices move slowly with long-term supply and demand; the price-spike era is over | "The Spot price of each instance type in each Availability Zone is set by Amazon EC2, and is adjusted gradually based on the long-term supply of and demand for Spot Instances." |