Inference & Serving

A Token Is Not a Unit: The Arithmetic Behind an Unforecastable AI Bill

One vendor shipped a new tokeniser that emits about 30 percent more tokens for the same text, at an unchanged published price. Nobody's rate card moved and everybody's bill rose. The industry buys AI by the token, and a token is not a unit of work, of text, or of anything stable enough to forecast against.

Anthropic's documentation carries a sentence that ought to unsettle anyone who forecasts an AI budget: models from Claude 4.7 onwards "use a newer tokenizer that contributes to their improved performance" and it "produces approximately 30% more tokens for the same text" (Anthropic, Pricing, consulted September 2026). The published price per million tokens did not rise. The rate card is unchanged, the model is better, and the same workload costs about a third more, because the thing being counted changed size.

This is not a scandal. It is a measurement problem, and it is the general case rather than an exception. Every organisation buying AI buys it in tokens, compares vendors in dollars per million tokens, and forecasts next quarter by multiplying an expected token volume by a price. Each of those three steps assumes a token is a unit: a fixed quantity of text, at a fixed price, representing a fixed amount of work. None of the three is true, and the gap between them and reality is where AI budgets go wrong, in a consistent direction.

Why this matters: A cost model built on dollars per million tokens cannot be diffed, cannot be compared across vendors, and drifts upward against actuals as a workload becomes more agentic. The replacement is not a better price list. It is measuring cost per completed task, on your own traffic, as a distribution rather than a mean.

TL;DR

  • A tokeniser change inside one model family can raise the cost of identical text by about 30 percent with no price change. Token counts describe the tokeniser, not the work.
  • One million input tokens on one model can bill at 0.05x, 0.1x, 1.0x, 1.25x or 2.0x of the list rate depending on which line item they land in, and the multipliers stack.
  • Agent loops re-send their whole context every turn, so billed input grows as \(O(n^2)\) in turns. Prompt caching divides the quadratic coefficient by exactly ten and leaves the exponent alone.
  • A 40-turn loop with an 8,000-token prefix bills 1,568,000 input tokens uncached and about 237,800 token-equivalents cached: a 6.6x cut that still grows quadratically.
  • Parts of the bill are not tokens: $10 per 1,000 searches, $0.05 per container-hour with a five-minute minimum, $0.08 per session-hour, 1.1x for pinned residency.
  • Tool definitions bill every turn. A 6,600-token toolset over 40 turns is 264,000 input tokens before the agent achieves anything.
  • In the worked triage model, caching cuts the token bill 2.2x and total cost per resolved ticket by 9 percent, because escalation dominates. Nine points of success rate are worth three times more.
  • Reasoning models emit roughly 18 times more tokens than non-reasoning counterparts and sometimes score lower, so token volume is partly the model's choice rather than yours.

At a Glance

flowchart LR
  W["One unit of work"] --> T1["Tokeniser<br/>counts it"]
  T1 --> T2["Loop re-sends<br/>the context"]
  T2 --> T3["Cache tier<br/>reprices it"]
  T3 --> T4["Non-token SKUs<br/>added"]
  T4 --> T5["Divided by<br/>success rate"]
  T5 --> B["What you<br/>actually pay"]
  classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
  classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
  classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
  class W blue
  class T1,T2,T3 purple
  class T4,T5 amber
  class B amber

Five transformations sit between a unit of work and a dollar figure. A price list describes one of them.

Why the Token Became the Price

Per-token pricing was a good idea for a defensible reason: in 2020 it tracked marginal cost closely. Prefill scales with input length because it is compute-bound; decode scales with output length because it is memory-bandwidth-bound. Charging separately for the two, at a ratio of roughly one to four, was an honest pass-through of two physical bottlenecks. For a single-turn completion on a dense model with no cache, dollars per million tokens was very nearly the right unit. Everything added since has widened the gap between the meter and the work.

timeline
    title How the unit came apart
    2020 : GPT-3 API ships per-token pricing
         : Dollars per thousand tokens becomes the comparison axis
    2023 : GPT-4 at 30 and 60 dollars per million sets the high-water mark
         : Input and output asymmetry becomes universal
    2024 : Prompt caching and asynchronous batch split one price into a menu
         : Epoch AI measures 9x to 900x annual declines per capability level
    2025 : Reasoning models bill hidden thinking at output rates
         : Agents measured at 4x chat tokens and multi-agent at 15x
    2026 : A tokeniser change adds about 30 percent tokens at an unchanged price
         : Server tools, container-hours and session-hours appear as non-token SKUs

The price fell hard over that period, and unevenly: Epoch AI found the price of reaching a fixed performance level dropping between 9 and 900 times per year depending on which level you pick (Cottier et al., 2025, LLM inference prices have fallen rapidly but unequally across tasks, Epoch AI). Volume rose faster. Google reported 9.7 trillion tokens a month two years earlier against 3.2 quadrillion by 2026, a more than 300-fold increase, with its model APIs at roughly 22 billion tokens a minute in the second quarter against 16 billion a quarter before (Pichai, 2026, Google I/O opening keynote; Pichai, 2026, Alphabet earnings call Q2 2026). Spending followed volume rather than price: Menlo Ventures put enterprise AI at $37 billion in 2025, 3.2 times the prior year (Menlo Ventures, 2025, The State of Generative AI in the Enterprise).

That divergence is usually told as a story about demand. It is also one about measurement: a unit that shrinks in price while the number of units per job grows, and whose definition quietly changes, is a unit you cannot plan with.

[IMAGE: Two-panel line chart on a shared time axis, 2023 to 2026. Left panel: price per million tokens for a fixed capability level, log scale, falling steeply. Right panel: tokens consumed per completed task for a representative workload, rising from single-call to agentic. Caption: "Both curves are real. Forecasts that use only the left one are the ones that miss."]

How the Unit Comes Apart

A token is not a fixed amount of text

The tokeniser decides how many tokens a string costs, and tokenisers change. Within one vendor's own lineup, the newer tokeniser emits about 30 percent more tokens for identical text, which raises cost and shrinks the usable context measured in characters, at an unchanged headline price. A migration plan that compares list prices across generations gets this exactly backwards.

[IMAGE: Side-by-side token-boundary visualisation of one identical paragraph under two tokenisers from the same model family, with boundaries drawn as vertical rules and the token count printed beneath each. The right panel shows roughly 30 percent more boundaries. Caption: "Same text, same price list, a third more tokens."]

The same effect runs across languages, where it is larger and has a fairness dimension. Evaluating ten models on 9,000 multiple-choice items across sixteen African languages, fertility measured in tokens per word reliably predicted accuracy, higher fertility predicting lower accuracy across every model and subject (Lundin et al., 2025, The Token Tax: Systematic Bias in Multilingual Tokenization, arXiv:2509.05486). The economic half of that finding is blunt: a doubling in tokens for the same meaning doubles the price of the same request, so one published price per token charges materially different amounts for identical work, and the customer does not control the difference. The tokenisation tax and token fertility across languages cover the mechanism; what matters here is that the billing unit inherits it.

A token is not a fixed price

The second decomposition is the price menu. The same input token, on the same model, is billed at a different rate depending on which of several line items it lands in (Anthropic, Pricing):

Line item Multiplier on base input What it buys
Fresh input 1.0x A full prefill
Five-minute cache write 1.25x A prefix reusable for five minutes
One-hour cache write 2.0x The same prefix for an hour
Cache hit or refresh 0.1x, or 0.05x on some models Skipping the prefill
Asynchronous batch 0.5x, input and output A 24-hour completion window
Pinned inference geography 1.1x on every category Data residency

The multipliers stack, so a cached read inside a batch job is 0.05x of base input and a residency-pinned version of it is 0.055x. The consequence is that a workload has an effective input rate rather than a price. If, over a billing period, a fraction \(\alpha\) of input tokens arrive fresh, \(\beta\) as cache writes at multiplier \(w\), and \(\gamma\) as cache reads at \(r\), then

\[p_{\text{eff}} = p_{\text{in}}\left(\alpha + w\beta + r\gamma\right), \qquad \alpha + \beta + \gamma = 1.\]

For a well-structured agent workload with \(\alpha = 0.05\), \(\beta = 0.10\), \(\gamma = 0.85\) at \(w = 1.25\) and \(r = 0.1\), the effective rate is \(0.05 + 0.125 + 0.085 = 0.26\) of list. For the same workload with an unstable prefix, where every request rewrites the cache, \(\beta \to 1\) and the effective rate is 1.25 of list. The published price is identical in both cases and the bills differ by a factor of nearly five. No procurement negotiation available to you moves the number that far.

[IMAGE: Stacked horizontal bar chart of effective input rate as a multiple of list price, for four prefix-stability regimes: no caching (1.0), unstable prefix (1.25), moderate hit rate (0.55), high hit rate (0.26). Each bar segmented into fresh, write and read contributions. Caption: "Prompt structure moves the effective rate further than model choice does."]

A token is not a unit of work

Here is the part that makes agentic budgets behave strangely. An agent re-sends its accumulated context on every turn, so the tokens billed are not the tokens of the conversation; they are the tokens of every prefix of the conversation, summed.

Let the stable prefix be \(P\) tokens of system instructions and tool definitions, and let each turn append \(a\) tokens of tool result and prior assistant output. Turn \(t\) sends \(P + (t-1)a\) input tokens, so over \(n\) turns

\[T_{\text{in}}(n) = \sum_{t=1}^{n}\bigl(P + (t-1)a\bigr) = Pn + a\,\frac{n(n-1)}{2}.\]

The quadratic term is the whole story. Doubling the turn count roughly quadruples the input bill, so a task that goes badly at 20 turns instead of 5 costs about sixteen times as much in input tokens, not four. Nothing in the price list hints at this, because the list is linear in tokens and the tokens are quadratic in turns.

Now add incremental caching, the way a real harness uses it. On turn \(t\) everything except the newest \(a\) tokens is already cached and reads at \(r = 0.1\), while the new \(a\) tokens are written at \(w = 1.25\); turn 1 writes the prefix. The billed token-equivalents are

\[T_{\text{eff}}(n) = wP + r P (n-1) + r\,a\,\frac{(n-2)(n-1)}{2} + w\,a\,(n-1).\]

Compare the quadratic coefficients: \(a/2\) uncached against \(ra/2 = a/20\) cached. Caching divides the quadratic coefficient by exactly ten and does not touch the exponent. It is a large constant-factor win on a growth law it cannot change, which is why "we turned on prompt caching" is a real saving and never a fix for a loop that runs long.

With \(P = 8{,}000\), \(a = 1{,}600\) and \(n = 40\):

  • Uncached: \(8{,}000 \times 40 + 1{,}600 \times 780 = 320{,}000 + 1{,}248{,}000 = 1{,}568{,}000\) input tokens. At $2 per million, $3.14.
  • Cached: \(10{,}000 + 31{,}200 + 118{,}560 + 78{,}000 = 237{,}760\) token-equivalents. At $2 per million, $0.48.

A 6.6x reduction, and the cached 40-turn loop still bills 2.6 times what a cached 20-turn loop would. The lever with the better exponent is turn count.

[IMAGE: Line chart of cumulative billed input tokens against turn number, 1 to 40, with two curves: uncached rising to 1.568 million and incrementally cached rising to about 238,000. Both visibly curve upward rather than straightening. Caption: "Caching moves the curve down by a factor of 6.6 and leaves its shape untouched."]

flowchart TB
  subgraph UC["Uncached loop"]
    U1["Turn 1: 8k"] --> U2["Turn 20: 38.4k"]
    U2 --> U3["Turn 40: 70.4k"]
    U3 --> UT["Total 1,568,000 tokens"]
  end
  subgraph CA["Incrementally cached loop"]
    C1["Turn 1: write 8k"] --> C2["Turn 20: read 36.8k, write 1.6k"]
    C2 --> C3["Turn 40: read 68.8k, write 1.6k"]
    C3 --> CT["Total 237,760 equivalents"]
  end
  UT --> R["Same O of n squared growth"]
  CT --> R
  classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
  classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
  classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
  class U1,U2,U3,UT rose
  class C1,C2,C3,CT emerald
  class R slate

A token is not the whole bill

Some of what you pay is not metered in tokens at all, and those terms grow with decisions rather than with text. Server-side search bills $10 per 1,000 searches however many results return. Code execution bills container time at $0.05 per hour beyond a monthly allowance of 1,550 hours, with a five-minute minimum per run, so two hundred three-second calls a day bill as roughly 16.7 hours rather than ten minutes. Managed agent sessions add $0.08 per session-hour, metered only while actually running. Pinning inference to one geography multiplies every token category by 1.1 (figures from Anthropic, Pricing, September 2026).

The subtlest term is still measured in tokens and still invisible to a per-request estimate: tool definitions are re-sent every turn. A browser toolset is documented at roughly 6,600 input tokens per request and a computer-use toolset at roughly 4,500, before any tool is called. Over a 40-turn loop that is 264,000 billed input tokens, $0.53 at $2 per million, for text the agent has already read 39 times. At the front of a cached prefix it costs about five cents; placed after something volatile, it is a line item nobody can find.

[IMAGE: Waterfall chart of one agent task's cost, starting at zero and adding bars for cached input, fresh input, output tokens, tool definitions, search fees and container time, with the non-token bars shaded amber. Caption: "Four of the six bars are invisible to a dollars-per-million-tokens estimate."]

A token is not a choice you control

Finally, the quantity of tokens a request consumes is partly the model's decision. Reasoning tokens bill at output rates, and their volume is set by the model's own behaviour under a requested effort level. Benchmarking 53 models across 14 basic arithmetic tasks, reasoning variants emitted roughly 18 times more tokens, sometimes at lower accuracy, with accuracy collapsing by up to about 36 percent when tokens were constrained and with non-monotonic accuracy-verbosity curves throughout (Srivastava et al., 2025, Do LLMs Overthink Basic Math Reasoning?, arXiv:2507.04023). Turning the effort dial up is therefore not reliably a purchase of accuracy: the harness behind the Holistic Agent Leaderboard found that higher reasoning effort reduced accuracy in the majority of its runs (Kapoor et al., 2026, Holistic Agent Leaderboard, ICLR, arXiv:2510.11977).

Fan-out is the same story at a larger grain. Anthropic measured agents using about four times the tokens of chat interactions and multi-agent systems about fifteen times, and found that token usage alone explained 80 percent of the performance variance on a browsing evaluation (Anthropic, 2025, How we built our multi-agent research system). Read that carefully: the variable driving quality is the variable driving spend. A token budget is a quality budget, and pretending otherwise is how cost programmes quietly ship regressions.

Seeing It in Motion

One turn of an agent loop touches four meters. Watching them tick separately is the fastest way to internalise why a single blended rate cannot describe the request.

sequenceDiagram
    participant A as Agent harness
    participant M as Model API
    participant T as Server tool
    A->>M: Send prefix plus history
    Note over M: Cached portion reads at 0.1x<br/>New delta writes at 1.25x
    M-->>A: Assistant turn plus tool call
    Note over M: Output and thinking bill at output rate
    A->>T: Execute search
    T-->>A: Results plus per-call fee
    Note over A,T: 10 dollars per 1,000 searches
    A->>M: Resend everything plus results
    Note over M: Context grew; next turn costs more

Four meters, three rates, one request. The usage object in each response is the only place all four appear together, which is why recording it per request is the difference between a cost model and a guess.

[IMAGE: Annotated screenshot-style mock of a single API response usage object, with callout arrows labelling input_tokens, cache_read_input_tokens, cache_creation_input_tokens, output_tokens and server_tool_use. Caption: "The five fields a cost model needs, and the one place they are reported together."]

Watch It Run

Animated diagram in which one unit of work flows left to right through tokenisation, an agent loop that visibly re-sends its growing context, a cache tier that reprices it, and non-token meters, before being divided by a success rate to produce the bill.
Solid animated edges are billed flows. The animated self-loop on the agent turn is the re-sent context, which is the quadratic term; the amber edges are the non-token meters that tick per decision rather than per token; the emerald edge is the cache read path that divides the quadratic coefficient by ten. The Mermaid figures above show the same structure if the animation is absent.

By the Numbers

Quantity Value Where it comes from
Cache read multiplier 0.1x of base input, 0.05x on some models Anthropic pricing, Sept 2026
Cache write multipliers 1.25x (5 min), 2.0x (1 hour) Anthropic pricing
Asynchronous batch discount 50% on input and output Anthropic pricing; matched by other major APIs
Pinned data residency 1.1x on every token category Anthropic pricing
Server-side search $10 per 1,000 searches Anthropic pricing
Code execution $0.05 per container-hour, 5-minute minimum, 1,550 free hours monthly Anthropic pricing
Managed agent session runtime $0.08 per session-hour, running time only Anthropic pricing
Browser toolset definition overhead about 6,600 input tokens per request Anthropic pricing
Tokeniser change within one family about 30% more tokens for the same text Anthropic pricing
Agent vs chat token usage about 4x; multi-agent about 15x Anthropic engineering blog
Share of BrowseComp variance from token usage 80% Anthropic engineering blog
Reasoning token multiplier about 18x, sometimes at lower accuracy Srivastava et al., arXiv:2507.04023
Annual price decline per capability level 9x to 900x Epoch AI
Google monthly token throughput 9.7T (2024) to 3.2 quadrillion (2026) Alphabet, I/O 2026
Enterprise AI market $37B in 2025, 3.2x year on year Menlo Ventures
Cost of validating one agent eval harness about $40,000 for 21,730 rollouts Kapoor et al., arXiv:2510.11977

Sources: multipliers and SKUs from Anthropic's published pricing as consulted in September 2026, which moves and should be re-checked rather than recalled. Token-usage multipliers from Anthropic's multi-agent write-up; reasoning figures from arXiv:2507.04023; price declines from Epoch AI; market size from Menlo Ventures; evaluation cost from arXiv:2510.11977. All vendor figures are list prices, not what a negotiated contract pays.

A Concrete Example

A support organisation handles 1 million tickets a month. A human touch costs $3.50 fully loaded. Three architectures are on the table, and the interesting result is not which one wins but by how much less than expected.

Shared assumptions: a 6,000-token prefix of system instructions, policy text and tool definitions; each agent turn appends 1,200 tokens of tool result and prior output and emits 350 tokens; two server-side searches per ticket at $10 per 1,000. Prices are the September 2026 list rates used above.

Step 1. Architecture A, one classification call on a small model. Input is the 6,000-token prefix plus a 900-token ticket; output 250 tokens. At $1 and $5 per million:

\[6{,}900 \times 10^{-6} \times 1 + 250 \times 10^{-6} \times 5 = \$0.0069 + \$0.00125 = \$0.00815.\]

Under a cent per ticket, and it resolves 38 percent of them without a person.

Step 2. Architecture B, an eight-turn agent, uncached, on a mid-tier model. Input tokens follow the loop law: \(6{,}000 \times 8 + 1{,}200 \times 28 = 48{,}000 + 33{,}600 = 81{,}600\). Output is \(8 \times 350 = 2{,}800\). At $2 and $10 per million, plus $0.02 of search fees:

\[81{,}600 \times 2\times10^{-6} + 2{,}800 \times 10\times10^{-6} + 0.02 = \$0.163 + \$0.028 + \$0.020 = \$0.211.\]

Twenty-six times the per-call cost of Architecture A. It resolves 71 percent.

Step 3. Architecture C, the same agent with an incrementally cached prefix. Using \(T_{\text{eff}}\) with \(P = 6{,}000\), \(a = 1{,}200\), \(n = 8\): \(7{,}500 + 4{,}200 + 2{,}520 + 10{,}500 = 24{,}720\) token-equivalents. So

\[24{,}720 \times 2\times10^{-6} + \$0.028 + \$0.020 = \$0.0494 + \$0.048 = \$0.0974.\]

The token bill fell by 2.2 times. Hold that number.

Step 4. Put the escalations back in. The cost that matters is per resolved ticket including the humans who handle the residue:

\[C_{\text{blended}} = C_{\text{model}} + (1 - s)\times \$3.50.\]
Architecture Model cost per ticket Success \(s\) Human cost per ticket Blended Monthly at 1M tickets
A, single call $0.0082 38% $2.170 $2.178 $2.18M
B, 8-turn agent, uncached $0.2112 71% $1.015 $1.226 $1.23M
C, same agent, cached $0.0974 71% $1.015 $1.112 $1.11M
C at 80% success $0.0974 80% $0.700 $0.797 $0.80M

Step 5. Read the table properly. Architecture A is 26 times cheaper per call and the most expensive option by a wide margin. Caching, which cut the token bill by 2.2x, cut the blended cost by 9 percent. Nine percentage points of extra success rate saved $0.315 per ticket, more than three times the entire caching win. At this operating point the token bill is not the lever, and a cost programme that optimises it is optimising a term worth a tenth of the one next to it.

That conclusion flips when escalation is cheap or absent. If the residue costs $0.20 rather than $3.50, the blended figures become $0.132, $0.269 and $0.155, Architecture A wins, and the token bill is suddenly most of the cost. The ranking of architectures is a function of the cost of being wrong, which is the formal result in Zellinger and Thomson, 2025, Economic Evaluation of LLMs, arXiv:2507.03834: reasoning models start paying once a mistake costs more than roughly a cent, and a single large model overtakes a cascade once a mistake costs about ten cents. Until you price an error, "cheaper" is not a well-formed claim.

[IMAGE: Grouped bar chart of blended cost per resolved ticket for architectures A, B, C and C-at-80-percent, each bar split into model cost and human escalation cost. A second small panel repeats the chart with escalation priced at 20 cents, showing the ranking invert. Caption: "The same three architectures, ranked twice, by the cost of being wrong."]

Where It Breaks

The effective rate is a moving average of your own behaviour

\(p_{\text{eff}}\) is not a price you were quoted; it is an outcome of prompt structure, traffic shape and hit rate, so it moves when nobody changed a contract. A refactor that inserts a per-request identifier above a stable policy block can take a workload from 0.26 of list to 1.25, a 4.8x increase, and it looks like a traffic anomaly on every dashboard because token counts barely move. Cache hit rate belongs next to latency and error rate in the service's own metrics, not in a monthly finance review.

A price list cannot be diffed

Two vendors publishing the same dollars per million tokens are not selling the same thing if their tokenisers, cache multipliers and free allowances differ, and one bills server-side tools per call while the other folds that capability into tokens. No arithmetic makes those comparable on paper. The only comparable quantity is measured cost per completed task on your own traffic, which is why a migration assessment is an engineering project rather than a spreadsheet. See cost per completed task.

Mean cost per request is the wrong forecasting input

Agentic cost distributions are heavy-tailed, because turn count enters quadratically, tool result sizes are chosen outside the system, and retries restart loops. If one percent of runs cost a hundred times the median and the rest cost the median, spend per run is \(0.99 + 1.00 = 1.99\) median-units and that one percent is 50.3 percent of the bill. The mean is then a tail artefact: it moves when rare runs get slightly longer. Forecast p50, p90 and p99 separately and treat p99/p50 as a leading indicator, since it moves before quality metrics do. See cost variance and the tail of agentic spend.

The unit changes under you

A tokeniser change, a new default reasoning effort, a deprecation that forces a migration, a toolset version whose definitions grew: each reprices existing traffic without any price changing. The defence is a pinned, versioned cost benchmark over a fixed set of representative requests, reporting tokens and dollars per completed task for the current model and harness. Without it you cannot separate a vendor-side change from your own regression, and both get attributed to whichever team shipped most recently.

Non-token meters resist attribution

A per-request token count can be attributed to a feature, a customer and a tenant. A container-hour with a five-minute minimum cannot be, cleanly: it was consumed by whichever run happened to fall inside the window. Free allowances make it worse by setting the apparent marginal price to zero until the month you cross the threshold, at which point the fitted model is wrong by the full list rate. Treat an allowance as a threshold with a forecast crossing date. See token accounting and cost attribution and non-token charges in an AI bill.

Cost-blind evaluation selects the expensive design

If your evaluation harness records accuracy and not cost, it will reliably prefer the more expensive architecture, because more tokens usually buy some accuracy. This was measured directly: on HumanEval a simple baseline matched a published state-of-the-art agent architecture at roughly 2 percent of its cost, a difference invisible on an accuracy-only leaderboard (Kapoor et al., 2024, AI Agents That Matter, arXiv:2407.01502). Recording cost per run is cheap. Not recording it means every architectural decision has a hidden thumb on the scale.

Alternative Designs

Design How it works Key advantage Key limitation Best when
List price per token Expected tokens times the rate card Trivial; vendor-published Blind to cache mix, non-token SKUs, tokeniser drift and success Single-turn, single-model, stable workloads
Effective rate per token Weight the rate card by the fresh, written and read mix Captures the largest structural discount Still linear in a quantity quadratic in turns Cache-heavy serving with stable prompts
Cost per completed task Measure tokens, fees and success on a golden set, then divide Comparable across vendors and architectures Needs a trusted rubric and an evaluation budget Any model or architecture choice
Cost per business outcome Put escalation and rework in the denominator The only figure a product owner can act on Requires owning the downstream process Features replacing a costed human workflow
Reserved capacity Buy throughput, not tokens Converts a variable bill into a fixed one Idle capacity is pure loss Steady base load at scale

Two of these are complements rather than alternatives. An effective-rate model is how you drive one design's cost down; cost per completed task is how you choose between designs. Running the first without the second is how teams arrive at a beautifully cached agent that loses to a two-line classifier. Owning the hardware instead is a separate calculation, worked in self-hosting versus API break-even.

[IMAGE: Decision flowchart. Start: "Which cost unit should I use?" Branches on whether the workload is single-turn, whether escalation is costed, whether you are choosing between models, and whether utilisation is flat, terminating in the six rows of the table above. Caption: "The unit follows the decision, not the invoice."]

How It Is Used in Practice

Public leaderboards moved first, because they had the strongest incentive. The SWE-bench leaderboards now carry per-instance cost next to resolution rate, which changes what a top position means. The Holistic Agent Leaderboard went further and made cost-controlled evaluation the infrastructure rather than an annotation, orchestrating 21,730 rollouts across nine models and nine benchmarks for about $40,000 (Kapoor et al., 2026, arXiv:2510.11977). Artificial Analysis reports a cost-to-run figure for its index built from each model's input, cache, reasoning and answer token prices and the tokens actually used, which is \(C_{\text{task}}\) with the success term set aside.

Inside product teams the practice that sticks is duller: record the full usage object on every request, tag it by feature, tenant and purpose, and keep a pinned benchmark of dollars per completed task that reruns on every model, prompt or harness change. That is what lets a team say whether the bill moved because traffic changed, because the workload changed, or because the vendor changed something. Teams that store only aggregate spend can answer none of the three.

The environmental disclosures now emerging are the same problem in another currency. Google reports a median Gemini text prompt at 0.24 Wh, 0.03 gCO2e and 0.26 mL of water, and its own narrower boundary, counting active accelerator draw alone, gives 0.10 Wh for the same prompt (Google Cloud, 2025, Measuring the environmental impact of AI inference). A factor of 2.4 from the accounting boundary, before any question of model or grid. Per-request energy is as slippery a unit as per-request cost, for the same reasons. See per-request energy, water and carbon.

Insights Worth Remembering

  1. The tokeniser is part of the price. A vendor can raise your bill by a third without changing a published number, and did. Cross-generation and cross-vendor comparisons have to measure token counts on your own text rather than inherit them.

  2. Prompt caching is a constant-factor win on a quadratic law. It divides the quadratic coefficient by ten and leaves the exponent alone: the highest-leverage change available without touching behaviour, and never a substitute for shortening the loop.

  3. Prefix stability is an interface with a price. One per-request value inserted above a stable block can multiply the effective input rate by nearly five. Treat prompt ordering as a contract and cache hit rate as a service metric.

  4. Tool definitions bill per turn, not per use. A 6,600-token toolset over 40 turns costs more than most whole requests. Load capability on demand and keep whatever must be declared inside the cached prefix.

  5. The variable driving quality is the variable driving spend. Token usage explained 80 percent of the performance variance on a browsing evaluation, so a token budget is a quality budget and a cap must be justified against completion rate.

  6. The ranking of architectures is a function of the cost of being wrong. Price the error first. Reasoning models, cascades and single large models each win in a different regime of that one parameter, and token-level optimisation does not move you between regimes.

  7. Optimise the largest term, not the most visible one. In the worked example a 2.2x cut in the token bill moved total cost by 9 percent, while nine points of success rate moved it by 28 percent. Token spend is the term with a dashboard, rarely the term with the money.

Open Questions

Can an outcome price be audited? Some vendors and applications now quote per resolved ticket rather than per token, which demonstrably moves tail risk from buyer to seller. Whether a buyer can verify an outcome claim without instrumenting the work themselves is unresolved, and the first generation of such contracts is likely to be priced from token models anyway, with margin covering the variance.

Does the quadratic loop law survive better context management? Compaction, external note-taking and code-execution tool interfaces all shrink what must be re-sent; one published workflow fell from 150,000 tokens to 2,000 by keeping results out of the window entirely. Whether these change the exponent or only the constants for realistic agents is open; the mechanism suggests they truncate the sum rather than flatten it.

What is the right unit for a heavy-tailed bill? Percentiles forecast well and price badly, because you cannot invoice a p99, and reserved capacity converts the problem into a utilisation problem rather than solving it. How a multi-tenant product should charge for a workload whose per-request cost spans two orders of magnitude is unsettled.

Will tokenisers converge or diverge? Fertility differences across languages are measured and consequential, and at least one proposal argues for fairer pricing rather than better modelling. Whether providers normalise against a language-independent unit or the divergence widens as tokenisers specialise is speculation, with commercial pressure pointing both ways.

Can cost-controlled evaluation hold as a norm? Recording cost per run is cheap and the evidence that omitting it distorts architecture choices is strong. Whether the practice survives the first time a cost column makes a favoured system look bad is sociological rather than technical, and the history of accuracy-only leaderboards is not encouraging.

Sources and Further Reading

  1. Anthropic. "Pricing." Claude Platform documentation, consulted September 2026. platform.claude.com/docs/en/about-claude/pricing
  2. Anthropic. "How we built our multi-agent research system." Anthropic Engineering, 2025. anthropic.com/engineering/multi-agent-research-system
  3. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). "AI Agents That Matter." arXiv:2407.01502
  4. Kapoor, S., Stroebl, B., Kirgis, P., et al. (2026). "Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation." ICLR 2026. arXiv:2510.11977
  5. Srivastava, G., Hussain, A., Srinivasan, S., & Wang, X. (2025). "Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models." arXiv:2507.04023
  6. Zellinger, M. J., & Thomson, M. (2025). "Economic Evaluation of LLMs." arXiv:2507.03834
  7. Wani, S. G., Dholakia, A., & Ellison, D. (2026). "The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts." arXiv:2608.26235
  8. Lundin, J. M., et al. (2025). "The Token Tax: Systematic Bias in Multilingual Tokenization." arXiv:2509.05486
  9. Bergemann, D., Bonatti, A., & Smolin, A. (2025). "The Economics of Large Language Models: Token Allocation, Fine-Tuning, and Optimal Pricing." Proceedings of the 26th ACM Conference on Economics and Computation (EC '25). doi:10.1145/3736252.3742625
  10. Cottier, B., et al. (2025). "LLM inference prices have fallen rapidly but unequally across tasks." Epoch AI. epoch.ai/data-insights/llm-inference-price-trends
  11. Menlo Ventures (2025). "2025: The State of Generative AI in the Enterprise." menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise
  12. Google Cloud (2025). "Measuring the environmental impact of AI inference." cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference
  13. Pichai, S. (2026). "Google I/O 2026: opening keynote." Google blog. blog.google/innovation-and-ai/sundar-pichai-io-2026
  14. Pichai, S. (2026). "Alphabet earnings call Q2 2026: Sundar Pichai remarks." Google blog. blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2026
  15. Anthropic (2025). "Code execution with MCP: building more efficient agents." Anthropic Engineering. anthropic.com/engineering/code-execution-with-mcp
  16. SWE-bench leaderboards, which report per-instance cost alongside resolution rate. swebench.com

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.