Spot & Interruptible Capacity
Deep discounts for work that can be interrupted and resumed.
5 to work through
-
intermediate
A data platform wants to use interruptible capacity to reduce cost. Which workloads are suitable, and what must be true of them?
2 min answer -
intermediate
A nightly batch job takes six hours on on-demand instances. Someone proposes spot to save 70%. What must be true?
2 min answer -
intermediate
A recommendation platform runs large offline training and feature computation jobs. What must be true architecturally before spot capacity is usable, and what does it cost?
2 min answer -
intermediate
Which workloads genuinely suit spot or preemptible capacity, what must be true of them, and what are the failure modes of over-applying it?
2 min answer -
advanced
A training pipeline runs on 400 interruptible GPU instances checkpointing every 30 minutes. Regional demand spikes and reclamation goes from a handful of instances an hour to most of the fleet inside ten minutes. What happens minute by minute, and what stops it?
2 min answer
3 terms in this topic
Interruption Tolerance
The property that determines whether a workload can use heavily discounted pre-emptible capacity, defined by what happens when an instance is reclaim…
conceptReclamation Notice Window
The seconds between a provider announcing it is taking an instance back and actually taking it, which sets a hard upper bound on how much work a shut…
practiceSpot Capacity
Using a provider's spare capacity at a steep discount, in exchange for the provider being able to reclaim it at short notice.
Neighbouring topics
Cost & FinOps
General material on the economics of an architecture.
Total Cost of Ownership
Lifetime cost including people, operations, upgrades and exit.
Cloud Pricing Models
On-demand, committed and spot, and the crossover arithmetic.
Reserved & Committed Capacity
Committing the baseline, laddering terms, and expiry as a silent failure.
Unit Economics
Cost per request, per tenant, per transaction — the actionable number.
Cost per Request
Fully-loaded per-unit cost, including the lines usually left out.
Storage Costs
Tiering, lifecycle policies, retrieval charges and minimum durations.
Network & Egress Costs
Cross-zone and cross-region transfer that appears on no diagram.
Compute Optimisation
Instance families, utilisation, and architecture that wastes less.
Rightsizing
Matching provisioned resources to observed demand rather than to a guess.
Cost Allocation
Tagging enforced at provisioning, and apportioning shared costs.
Showback & Chargeback
Making teams see, or own, the cost of what they run.
Architecture Cost Modelling
Pricing a design before building it, at expected and at ten times volume.
Cost vs Reliability
Each nine costing an order of magnitude, and pricing the failure instead.
Observability Cost
Telemetry bills, cardinality control and retention tiering.
Licence & Vendor Costs
Per-core, per-seat and per-environment terms that shape designs.
FinOps Practice
Inform, optimise, operate — and cost as a design-review criterion.
Waste Elimination
Idle environments, orphaned volumes, forgotten instances, dead data.
Cost Governance
Budgets, anomaly alerts, quotas and preventive policy.