Capacity Reservation
also called On-Demand Capacity Reservation, Zonal Capacity Hold
Paying to hold specific hardware in a specific zone whether or not it is running, which is the only purchase that guarantees a launch succeeds and is separate from every discount instrument.
A retail platform holds a three-year spend commitment covering 85% of its fleet and treats it as capacity planning. On the biggest trading day of the year the autoscaler asks for 200 more instances of its usual type in its usual zone and receives insufficient-capacity errors. The commitment keeps discounting the instances that are already running and does nothing at all about the ones that will not start.
A discount instrument and a capacity guarantee are different products with opposite billing behaviour. Spend commitments and regional reserved instances lower the rate on usage that occurs; AWS documents that Savings Plans do not reserve capacity and that regional reserved instances reserve none either. A capacity reservation holds hardware in one availability zone for one instance type, bills at the on-demand rate for every hour it exists whether or not anything runs in it, and carries no discount of its own. The provider's guidance for business-critical workloads is to buy both: the reservation for the machine, the commitment for the price, with the discount applying to the reservation's charges.
Why it matters
Cloud capacity is finite per zone per instance type, and the moments when it runs short are correlated with the moments everyone wants it: a retail peak, a regional failover, a popular accelerator generation. The failure mode is not a slow system but a system that cannot grow, and it arrives as an API error in an autoscaling group rather than as anything a cost dashboard shows.
The second reason is financial. A commitment is charged per hour regardless of usage, so a capacity shortage means paying for committed spend you could not consume. The shortage and the waste arrive together.
Implementation patterns
- Reserve only the slice whose absence is an incident: the peak-hour baseline of the revenue path, not the whole fleet. 10% of a fleet held for a two-week peak costs roughly that fraction of on-demand price for the period.
- Create reservations days before a known event and release them after, so the on-demand-rate hours are bounded. Treat the create and release as steps in the event runbook.
- Let the discount apply on top. Reservation charges are usage like any other and are covered by a commitment, so the combination costs the committed rate rather than list.
- Diversify instead, where you can. An autoscaling group spread over six instance types and three zones survives a single-pool shortage without paying for idle hardware.
- Reserve for failover arithmetic, not for steady state: a three-zone service that must survive losing one needs the surviving zones to be able to launch, which is a capacity question nobody tests.
Industry example
The accelerator market made this distinction visible to everyone. When a GPU generation is supply constrained, a spend commitment buys nothing a scheduler can use, and organisations training large models negotiate explicit capacity - reservations, dedicated blocks, or multi-year contracts naming hardware and dates. The economics invert too: an idle reserved accelerator costs more per hour than most teams' entire compute bill, so reservation sizing becomes a utilisation problem, solved with a queue that keeps the reserved block busy with lower-priority work between the jobs it was bought for.
Failure scenarios
- Commitment mistaken for capacity, discovered at peak, with no time to fix.
- Reservations created and never released, billing on-demand rate for months after the event.
- A failover that cannot land because the surviving zones have no capacity for the displaced workload.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Capacity reservation | a launch that is guaranteed to succeed | on-demand rate for 100% of hours held |
| Diversified allocation | no idle cost | no guarantee in a broad shortage |
| Zonal reserved instance | discount and capacity together | zone lock-in for the whole term |
When not to use it
For a workload that is genuinely family-agnostic and zone-flexible, a diversified allocation strategy delivers the same protection for nothing, and a paid reservation is an expensive answer to a solved problem. For workloads that can wait - batch, training, reprocessing - queueing is cheaper than reserving. And a reservation is never justified by comfort: it is justified by the cost of the outage it prevents, which is a number somebody has to produce.
Interview question
Q: Your CFO has signed a three-year commitment covering 85% of compute and considers capacity planning done. The retail peak is in six weeks. What do you tell them, and what do you actually buy?
What a strong answer covers: the price-versus-capacity distinction with its billing consequences; which slice of the fleet needs a guarantee and why the rest does not; the create-and-release runbook; the fact that commitments are charged even when usage falls short, so a shortage wastes money twice; and a rehearsal date before the event.
Quick check
Quiz: Does a Savings Plan guarantee you can launch an instance? No. It discounts usage that happens; only a zonal reserved instance or a capacity reservation holds hardware, and the latter bills whether used or not.
Flashcard: What does a capacity reservation cost when nothing runs in it? — The full on-demand rate for every hour it exists, which is why it is created before an event and released after.