The Price Menu: Caching, Batch and Service Tiers
The multipliers that sit between a model's list price and what you actually pay, the break-even arithmetic for a cache write, and why the menu exists as price discrimination rather than generosity.
The same million input tokens, on the same model, can be billed at twenty times different rates depending on which line item they land in. Read from a warm prompt cache they cost a tenth of the base rate; written to a one-hour cache they cost twice it. Sent through the asynchronous batch endpoint they are halved. Pinned to a single inference geography they carry a further ten percent. None of this appears in the headline number a procurement spreadsheet copies.
The multipliers
Anthropic publishes the multipliers explicitly, which makes them usable as arithmetic rather than folklore (Anthropic, Pricing, figures as of September 2026):
| Line item | Multiplier on base input price | What you are buying |
|---|---|---|
| Base input | 1.0x | A fresh prefill |
| Five-minute cache write | 1.25x | A prefix reusable for five minutes |
| One-hour cache write | 2.0x | The same prefix, reusable for an hour |
| Cache hit or refresh | 0.1x | Skipping the prefill entirely |
| Batch endpoint | 0.5x on input and output | A 24-hour completion window |
| Fast mode (where offered) | 2.0x | Higher output speed |
| Pinned inference geography | 1.1x on every category | Data residency |
Two properties matter more than the individual values. The multipliers stack, so a cached read inside a batch job is 0.05x of base input. And they apply per token, not per request, so the mix of cached, fresh and output tokens decides the effective rate far more than the model choice does. Some models depart from the table: a cache hit on Claude Opus 5.5 is 0.05x rather than 0.1x, so two models in one family can rank differently on a cache-heavy workload than their list prices suggest.
Cache break-even, derived
Caching is not free, so it is worth knowing when it pays. Let a prefix of \(L\) tokens be reused across \(n\) total requests, at base input price \(p\), with a write multiplier \(w\) and a read multiplier \(r = 0.1\). Without caching you pay \(nLp\). With caching you pay one write and \(n-1\) reads:
Caching wins when \(w + r(n-1) < n\), that is when
For the five-minute cache, \(w = 1.25\) gives \(n > 1.28\), so two requests, one write and one read, already pay for it. For the one-hour cache, \(w = 2.0\) gives \(n > 2.11\), so three requests are needed. This matches the vendor's own guidance and, more usefully, shows what the longer time-to-live is actually for: it does not buy a better price, it buys a wider window in which the second request may arrive. On a low-traffic endpoint where requests are minutes apart, the 2x write is the cheaper choice precisely because the 1.25x write would expire unread.
The variable that dominates in practice is not \(w\) but the hit rate. A prefix that changes on every request never reads, so every write is pure overhead: a workload with a 0 percent hit rate and full caching enabled pays 1.25x for nothing.
Why the menu exists
These are not discounts, they are a screening mechanism. A provider serving one model to buyers with wildly different latency tolerance and volume can extract more surplus, and fill more of its capacity trough, by offering a menu than by setting one price. The formal result is that under high-dimensional user heterogeneity the optimal mechanism is implementable as menus of two-part tariffs, with higher markups for more intensive users, which is close to what the published price lists look like (Bergemann, Bonatti and Smolin, 2025, The Economics of Large Language Models: Token Allocation, Fine-Tuning, and Optimal Pricing, ACM EC'25).
Read that way, the batch tier is the provider buying schedule flexibility from you at a 50 percent discount, for the same reason a utility sells interruptible load cheaply. The mirror of capacity planning under lumpy demand is that your deferrable work is the thing that smooths someone else's fleet. Competing providers price the same two mechanisms almost identically: a 50 percent asynchronous batch discount and a 90 percent discount on cached input are now common across major APIs, alongside explicit latency tiers that charge a premium for priority scheduling.
When it breaks
A single unstable byte destroys the hit rate. Cache reads require an exact prefix match, so a timestamp, a request ID or a shuffled tool list at the top of the prompt converts every read into a write. Order the prompt from most stable to least, and treat prefix stability as an interface with a cost attached.
Writes are charged whether or not they are read. The break-even above assumes the reuse happens. Enabling caching on a long prefix for traffic that never repeats is a 25 percent input-price increase.
Batch is not a latency knob. A 24-hour window is unusable for anything interactive, and a job that must complete before a downstream deadline cannot be batched even if it is offline in principle. See batch and asynchronous LLM workloads.
Effective rate hides model substitution. A workload re-engineered to a 90 percent cache hit rate on an expensive model can beat the same workload uncached on a cheap one, so model choice and prompt structure cannot be optimised separately.
Compaction and caching fight. Rewriting a conversation to shorten it produces a new prefix, so the next request pays a full prefill at 1.0x or a write at 1.25x. The token saving is real and the cache saving is forfeited.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.