Cost & FinOps for AI intermediate 7 min read 12 flashcards

Reserved, On-Demand and the Shape of Commitment

How to choose a commitment mix when demand is uncertain, why the break-even is simply a price ratio, and the option value that makes shorter commitments rational despite costing more.

Every capacity decision is a bet on future demand. Commit too much and you pay for idle hardware; commit too little and you pay a premium for the difference, or cannot get it at all. The arithmetic is simple and the inputs are uncertain, so the useful skill is reasoning about the uncertainty rather than computing the break-even.

The break-even

If committed capacity costs a fraction \(r\) of the on-demand rate, it pays off when utilisation exceeds \(r\). At 60 percent of on-demand, commit for the capacity you will use more than 60 percent of the time. This is exact and it assumes the two are interchangeable, which for scarce accelerators they are not.

The correct structure follows from the demand distribution rather than from its average. Commit to the level you are confident of using continuously, which is the baseline of predictable serving and scheduled work. Serve the variable portion on demand or interruptibly. Committing to the mean guarantees paying for idle capacity roughly half the time.

Option value

Shorter commitments cost more per hour and buy the right to change your mind. That right has value in proportion to the uncertainty: a new accelerator generation may arrive, demand may fall, an architectural change may cut inference cost by half, or the workload may move to a different provider.

This is why a three-year commitment at a deep discount is not automatically better than a one-year commitment at a smaller one. The comparison is between the discount gained and the option surrendered, and in a field where price-performance has improved rapidly and repeatedly, the option is worth more than the discount arithmetic alone suggests.

The self-hosting comparison

Running open-weights models on owned or reserved hardware converts variable cost to fixed. It is cheaper above a utilisation threshold that is often surprisingly high once the full picture is included: hardware or reservation, engineering time to build and operate the serving stack, on-call burden, and the capacity headroom needed for peaks.

The comparison is also not like-for-like on capability. A hosted frontier model and a self-hosted open-weights model are different products, and a cost comparison that ignores the quality difference is comparing the prices of two things without noting they are not the same thing.

When it breaks

Utilisation forecasts are optimistic. Teams reliably overestimate how fully they will use reserved capacity, because the plan assumes the workloads that were promised arrive on time. Committing to a level below the forecast, and topping up on demand, is the asymmetric choice: the cost of under-committing is a premium rate, the cost of over-committing is the whole amount.

Commitments are not always transferable. Some can be resold, exchanged or applied flexibly across instance types; others cannot. The flexibility terms are a substantial part of the value and are frequently not read until they matter.

Generational transitions strand commitments. A commitment on the current generation becomes expensive relative to the next one's price-performance, and the transition is not a surprise: cadences are roughly known. Aligning commitment terms with that cadence is a straightforward improvement that is rarely made deliberately.

Availability trumps price for scarce parts. When a specific accelerator type is constrained, a commitment buys the ability to run at all. Under those conditions the break-even calculation is beside the point, and the decision is about securing capacity rather than optimising its rate.

Check yourself

12 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track