advanced 3 min answer

Your product puts a 30000-token policy document in the system prompt of every model request and serves 40 tenants through one shared prompt cache. Anthropic's documentation in 2026 prices cache writes at 1.25x the base input rate and cache reads at 0.1x for most models. Design the per-tenant cost allocation.

anthropicprompt-cachingcost-allocationshared-costmulti-tenant
Show the full answer Hide the answer

Requirements that drive the structure

The cost of a request depends on whether its prefix was already cached, which depends on who sent a request before it. Any allocation scheme that charges each tenant the rate its own request happened to pay will produce bills that differ several-fold for identical work, so the scheme has to separate the shared prefix from the tenant's own tokens.

The arithmetic that sets the design

Express everything in base-input-token equivalents for the 30000-token prefix:

  • A cache write costs 30000 × 1.25 = 37500 equivalents.
  • A cache read costs 30000 × 0.1 = 3000 equivalents.

So the first request after a miss pays 12.5x what the next one pays, for the same prefix. Amortise across the requests that share one write inside the cache lifetime:

  • 400 requests share it: (37500 + 399 × 3000) / 400 ≈ 3086 equivalents each, within 3% of the read price.
  • 2 requests share it: (37500 + 3000) / 2 = 20250 each, nearly 7x more.

The marginal cost of a request is therefore a property of the traffic pattern, not of the request.

The design

  1. Two cost pools per model. A shared pool holding every cache write, and a direct pool holding each tenant's own variable input tokens and output tokens, which are unambiguously theirs.
  2. Allocate the shared pool by reads in the same window. Divide each write across the requests that read that prefix within the cache lifetime, and never charge it to the tenant that triggered it.
  3. Report both lines. Direct cost and allocated share, with the key printed, so a tenant disputing its bill can recompute it.
  4. Instrument the signal, not the estimate. Log cache read and cache creation token counts per request from the API response, so the allocation is measured rather than modelled.

The hard part

Attribution across a time window, not a request. A write is consumed by reads until the entry expires, so the denominator is only known once the window closes. Batch the allocation per window and accept a short reporting lag rather than guessing in real time.

What it costs

Engineering: a usage table keyed by tenant, model and window, and a nightly job. Perhaps two engineer-weeks. Prefer a flat per-request rate unless total model spend is above roughly $5k a month, because below that the machinery costs more than the allocation error it removes.

How it fails

  • Charging the write to the first requester. The low-volume tenant pays up to 12.5x for identical work, and a team shown that bill will add a keep-warm pinger, which spends more than the write it avoids.
  • A volatile prefix. Any byte change invalidates the cached prefix, so a timestamp or a request id in the system prompt turns every request into a write: 37500 instead of 3000 equivalents, a 12.5x cost increase with no visible symptom. The signal is cache read tokens sitting at zero across repeated requests.
  • A prefix below the minimum. Anthropic's documented minimum cacheable prefix in 2026 ranges from 512 to 4096 tokens by model, and shorter prefixes are processed without caching and without an error, so the model predicts a discount the invoice never shows.

What I would build first

The per-request usage log with cache read and write token counts. Everything else is arithmetic over that table, and without it both the allocation and the caching decision are guesses.

When this is the wrong answer

If the prefix differs per tenant there is no shared cost to allocate: each tenant has its own entry, pays its own write, and a direct per-tenant sum is simpler and exactly right. The design above exists only because the expensive thing is shared. The same applies under a flat subscription with no usage component, where the allocation changes no decision anyone makes.