practice

Prompt Cache Economics

also called Prefix Cache Break-Even, Cache TTL Selection

Deciding what to cache and at which time-to-live by comparing the provider's write premium against its read discount, so a long stable prefix is paid for once rather than on every request.

ai-cost-managementprompt-cachingttlprefix-stabilityunit-economics

An assistant sends a 12,000-token prefix on every request: system prompt, tool catalogue, policy document, few-shot block. The question adds 40 tokens. Without caching, the bill is dominated by text that has not changed in a month.

Providers price this directly. Anthropic's published rates in 2026: writing a prefix into the cache costs 1.25x the base input rate at a 5-minute time-to-live and 2x at a 1-hour one, while reading it costs 0.1x base for most models and 0.05x on Claude Opus 5.5. The minimum cacheable prefix is model-dependent, between 512 and 4,096 tokens, and anything shorter is silently not cached. These multipliers change, so carry them as configuration rather than as knowledge.

Why it matters

The break-even is startlingly short, which is why this is the first cost lever to pull rather than the last.

Five-minute TTL. Two uncached calls cost 2.0 units of input. Cached: 1.25 to write plus 0.1 to read — 1.35 units, so the first reuse already wins; a tenth reuse costs 2.15 against 10.0.

One-hour TTL. A 2.0 write plus a 0.1 read is 2.1 against 2.0 uncached: break-even at the second read and positive from the third.

Latency improves for free, because a cached prefix is not processed again, and the gain lands on the long-context requests that feel slowest.

Implementation patterns

  • Freeze the prefix. Caching is a byte-exact prefix match, rendered as tools, then system, then messages, so a timestamp, request id, user name or unsorted JSON dump in that region invalidates everything after it.
  • Order by volatility: stable instructions first, the breakpoint next, volatile content after it, with the tool list in a deterministic order.
  • Pick the TTL from the reuse interval. Inside a conversation or a burst: 5 minutes. A few times an hour: 1 hour. Less often than the TTL: do not cache, since the write premium is paid every time.
  • Verify, do not assume. Read the cache-read token counter; zero reads across repeated identical prefixes means a silent invalidator.
  • Pre-warm before a known burst, so a launch does not pay thousands of write premiums.

Industry example

The published multipliers are the concrete case, documented because the behaviour is counter-intuitive: a discount that charges more than base rate on first use. The minimum-prefix rule matters as much — caching a 300-token system prompt does nothing and reports no error, which is how teams conclude caching "did not work".

Failure scenarios

  • One line destroys the discount. A current timestamp appended to the system prompt so the model knows the date drops the hit rate to near zero. Latency and errors barely move; the bill multiplies.
  • A cache-miss storm after deploy. A changed system prompt makes every active session pay the write premium at once — a cost spike with no error signal.
  • A mid-conversation parameter change — another model, a changed tool set, a top-level effort change — invalidates the prefix, so a feature toggle becomes a cost regression.
  • An over-long TTL on a rarely used prompt, where the 2x write is paid repeatedly and caching is strictly worse than not caching.

Trade-offs

Choose Gains Pays
No caching Nothing to reason about Full input price on every repeated token
5-minute TTL Wins from the first reuse A 1.25x premium on content used only once
1-hour TTL Covers gaps between bursts 2x on write, so it needs two or more reads
Pre-warming Protects a launch from a write storm An extra call and the discipline to run it

When not to use it

Do not cache when the prefix is not reused inside the TTL. A low-traffic internal tool called twice a day pays the write premium each time and saves nothing, and short one-shot classification has no shared prefix to amortise.

Do not contort the prompt to make caching work either. If personalising the system prompt measurably improves quality, keep it and move it after the breakpoint, so shared instructions still cache and only the personal block is fresh. The discount is worth a lot; it is not worth a worse answer.

Interview question

Q: Our input bill tripled overnight. No model change, no traffic change, latency unchanged, error rate unchanged. Where do you look, and what would have caught it?

What a strong answer covers: cache read and write counters first, because a hit-rate collapse is the only mechanism that triples input cost with nothing else moving; then the prefix-rendering diff for a timestamp, user field, unsorted serialisation or reordered tool list; the deploy-time write storm as a second candidate; moving volatile content after the breakpoint; and the monitoring that makes it visible, a cache hit-rate SLI with an alert threshold and cost per interaction tracked per release.

Quick check

Quiz: A 5-minute write costs 1.25x base input and a read 0.1x. After how many uses does caching pay? — The second (1.35 against 2.0), which is why a stable prefix is cached by default rather than after an analysis.

Flashcard: When is a 1-hour TTL the wrong choice? — When the prefix is reused less than twice an hour: the 2x write makes a single read more expensive than not caching. Match the TTL to the measured reuse interval.