A travel marketplace of Expedia's kind holds a three-year compute commitment covering 85% of its steady fleet. During a regional peak an availability zone runs short of its instance type and the autoscaler cannot launch. What happens to the bill and to the capacity?
Show the full answer Hide the answer
What is being tested
The distinction between buying a price and buying a machine. Finance language merges them. The provider sells them as separate products with opposite billing behaviour, and the confusion surfaces on the one day capacity matters.
What happens, in order
A Savings Plan or a regional reserved instance is a billing construct. AWS documents it plainly: Savings Plans do not reserve capacity, and regional reserved instances have no capacity reserved either. Only a zonal reserved instance or an On-Demand Capacity Reservation holds hardware in a zone, and a capacity reservation carries no discount while billing whether you use it or not. AWS's own guidance for business-critical workloads is to hold a capacity reservation for the capacity and a Savings Plan for the price.
So, minute by minute: the autoscaler gets insufficient-capacity errors and retries. Running instances keep their discount. The hourly commitment is charged in full whether or not usage reaches it, so a capacity shortfall silently converts committed spend into waste and the effective discount rate falls at the worst moment. The fleet is short of capacity and the bill is unchanged.
Why the other options fail
- Priority over on-demand customers. The most expensive misreading in the set, because it is how a reservation sounds. Regional commitments confer no scheduling priority of any kind.
- Rolling unused commitment forward. Comforting and false. Commitments are consumed per hour; an hour of under-use is gone, which is why coverage is sized to the trough rather than the average.
- Converting to on-demand. A commitment is an obligation, not a credit balance. There is nothing to convert, and the obligation is exactly what makes under-use expensive.
The design that would have survived it
Split the question in two. Reserve capacity for the slice of the fleet whose absence is an incident - typically the peak-hour baseline of the checkout and search paths - with explicit capacity reservations created days before the event and released after. Buy the discount separately, sized to the trough. Diversify instance families and zones in the autoscaling group so a shortage in one pool is survivable. Then rehearse: launch the peak shape a week early, because an insufficient-capacity error during a rehearsal is information and the same error at peak is revenue.
When this is the wrong answer
If the workload is genuinely family-agnostic and zone-flexible, allocation strategies across six instance types solve this for free and paid capacity reservations are an expensive answer to a solved problem. Holding 10% of a fleet as capacity reservations for a two-week peak costs roughly that fraction of on-demand price for the whole period, so it is justified by the cost of the outage it prevents and by nothing else.