A Token Is Not a Unit: The Arithmetic Behind an Unforecastable AI Bill
One vendor shipped a new tokeniser that emits about 30 percent more tokens for the same text, at an unchanged published price. Nobody's rate card moved and everybody's bill rose. The industry buys AI by the token, and a token is not a unit of work, of text, or of anything stable enough to forecast against.
Anthropic's documentation carries a sentence that ought to unsettle anyone who forecasts an AI budget: models from Claude 4.7 onwards "use a newer tokenizer that contributes to their improved performance" and it "produces approximately 30% more tokens for the same text" (Anthropic, Pricing, consulted September 2026). The published price per million tokens did not rise. The rate card is unchanged, the model is better, and the same workload costs about a third more, because the thing being counted changed size.
This is not a scandal. It is a measurement problem, and it is the general case rather than an exception. Every organisation buying AI buys it in tokens, compares vendors in dollars per million tokens, and forecasts next quarter by multiplying an expected token volume by a price. Each of those three steps assumes a token is a unit: a fixed quantity of text, at a fixed price, representing a fixed amount of work. None of the three is true, and the gap between them and reality is where AI budgets go wrong, in a consistent direction.
Why this matters: A cost model built on dollars per million tokens cannot be diffed, cannot be compared across vendors, and drifts upward against actuals as a workload becomes more agentic. The replacement is not a better price list. It is measuring cost per completed task, on your own traffic, as a distribution rather than a mean.
TL;DR
- A tokeniser change inside one model family can raise the cost of identical text by about 30 percent with no price change. Token counts describe the tokeniser, not the work.
- One million input tokens on one model can bill at 0.05x, 0.1x, 1.0x, 1.25x or 2.0x of the list rate depending on which line item they land in, and the multipliers stack.
- Agent loops re-send their whole context every turn, so billed input grows as \(O(n^2)\) in turns. Prompt caching divides the quadratic coefficient by exactly ten and leaves the exponent alone.
- A 40-turn loop with an 8,000-token prefix bills 1,568,000 input tokens uncached and about 237,800 token-equivalents cached: a 6.6x cut that still grows quadratically.
- Parts of the bill are not tokens: $10 per 1,000 searches, $0.05 per container-hour with a five-minute minimum, $0.08 per session-hour, 1.1x for pinned residency.
- Tool definitions bill every turn. A 6,600-token toolset over 40 turns is 264,000 input tokens before the agent achieves anything.
- In the worked triage model, caching cuts the token bill 2.2x and total cost per resolved ticket by 9 percent, because escalation dominates. Nine points of success rate are worth three times more.
- Reasoning models emit roughly 18 times more tokens than non-reasoning counterparts and sometimes score lower, so token volume is partly the model's choice rather than yours.
At a Glance
flowchart LR W["One unit of work"] --> T1["Tokeniser<br/>counts it"] T1 --> T2["Loop re-sends<br/>the context"] T2 --> T3["Cache tier<br/>reprices it"] T3 --> T4["Non-token SKUs<br/>added"] T4 --> T5["Divided by<br/>success rate"] T5 --> B["What you<br/>actually pay"] classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff class W blue class T1,T2,T3 purple class T4,T5 amber class B amber
Five transformations sit between a unit of work and a dollar figure. A price list describes one of them.
Why the Token Became the Price
Per-token pricing was a good idea for a defensible reason: in 2020 it tracked marginal cost closely. Prefill scales with input length because it is compute-bound; decode scales with output length because it is memory-bandwidth-bound. Charging separately for the two, at a ratio of roughly one to four, was an honest pass-through of two physical bottlenecks. For a single-turn completion on a dense model with no cache, dollars per million tokens was very nearly the right unit. Everything added since has widened the gap between the meter and the work.
timeline
title How the unit came apart
2020 : GPT-3 API ships per-token pricing
: Dollars per thousand tokens becomes the comparison axis
2023 : GPT-4 at 30 and 60 dollars per million sets the high-water mark
: Input and output asymmetry becomes universal
2024 : Prompt caching and asynchronous batch split one price into a menu
: Epoch AI measures 9x to 900x annual declines per capability level
2025 : Reasoning models bill hidden thinking at output rates
: Agents measured at 4x chat tokens and multi-agent at 15x
2026 : A tokeniser change adds about 30 percent tokens at an unchanged price
: Server tools, container-hours and session-hours appear as non-token SKUsThe price fell hard over that period, and unevenly: Epoch AI found the price of reaching a fixed performance level dropping between 9 and 900 times per year depending on which level you pick (Cottier et al., 2025, LLM inference prices have fallen rapidly but unequally across tasks, Epoch AI). Volume rose faster. Google reported 9.7 trillion tokens a month two years earlier against 3.2 quadrillion by 2026, a more than 300-fold increase, with its model APIs at roughly 22 billion tokens a minute in the second quarter against 16 billion a quarter before (Pichai, 2026, Google I/O opening keynote; Pichai, 2026, Alphabet earnings call Q2 2026). Spending followed volume rather than price: Menlo Ventures put enterprise AI at $37 billion in 2025, 3.2 times the prior year (Menlo Ventures, 2025, The State of Generative AI in the Enterprise).
That divergence is usually told as a story about demand. It is also one about measurement: a unit that shrinks in price while the number of units per job grows, and whose definition quietly changes, is a unit you cannot plan with.
[IMAGE: Two-panel line chart on a shared time axis, 2023 to 2026. Left panel: price per million tokens for a fixed capability level, log scale, falling steeply. Right panel: tokens consumed per completed task for a representative workload, rising from single-call to agentic. Caption: "Both curves are real. Forecasts that use only the left one are the ones that miss."]
How the Unit Comes Apart
A token is not a fixed amount of text
The tokeniser decides how many tokens a string costs, and tokenisers change. Within one vendor's own lineup, the newer tokeniser emits about 30 percent more tokens for identical text, which raises cost and shrinks the usable context measured in characters, at an unchanged headline price. A migration plan that compares list prices across generations gets this exactly backwards.
[IMAGE: Side-by-side token-boundary visualisation of one identical paragraph under two tokenisers from the same model family, with boundaries drawn as vertical rules and the token count printed beneath each. The right panel shows roughly 30 percent more boundaries. Caption: "Same text, same price list, a third more tokens."]
The same effect runs across languages, where it is larger and has a fairness dimension. Evaluating ten models on 9,000 multiple-choice items across sixteen African languages, fertility measured in tokens per word reliably predicted accuracy, higher fertility predicting lower accuracy across every model and subject (Lundin et al., 2025, The Token Tax: Systematic Bias in Multilingual Tokenization, arXiv:2509.05486). The economic half of that finding is blunt: a doubling in tokens for the same meaning doubles the price of the same request, so one published price per token charges materially different amounts for identical work, and the customer does not control the difference. The tokenisation tax and token fertility across languages cover the mechanism; what matters here is that the billing unit inherits it.
A token is not a fixed price
The second decomposition is the price menu. The same input token, on the same model, is billed at a different rate depending on which of several line items it lands in (Anthropic, Pricing):
| Line item | Multiplier on base input | What it buys |
|---|---|---|
| Fresh input | 1.0x | A full prefill |
| Five-minute cache write | 1.25x | A prefix reusable for five minutes |
| One-hour cache write | 2.0x | The same prefix for an hour |
| Cache hit or refresh | 0.1x, or 0.05x on some models | Skipping the prefill |
| Asynchronous batch | 0.5x, input and output | A 24-hour completion window |
| Pinned inference geography | 1.1x on every category | Data residency |
The multipliers stack, so a cached read inside a batch job is 0.05x of base input and a residency-pinned version of it is 0.055x. The consequence is that a workload has an effective input rate rather than a price. If, over a billing period, a fraction \(\alpha\) of input tokens arrive fresh, \(\beta\) as cache writes at multiplier \(w\), and \(\gamma\) as cache reads at \(r\), then
For a well-structured agent workload with \(\alpha = 0.05\), \(\beta = 0.10\), \(\gamma = 0.85\) at \(w = 1.25\) and \(r = 0.1\), the effective rate is \(0.05 + 0.125 + 0.085 = 0.26\) of list. For the same workload with an unstable prefix, where every request rewrites the cache, \(\beta \to 1\) and the effective rate is 1.25 of list. The published price is identical in both cases and the bills differ by a factor of nearly five. No procurement negotiation available to you moves the number that far.
[IMAGE: Stacked horizontal bar chart of effective input rate as a multiple of list price, for four prefix-stability regimes: no caching (1.0), unstable prefix (1.25), moderate hit rate (0.55), high hit rate (0.26). Each bar segmented into fresh, write and read contributions. Caption: "Prompt structure moves the effective rate further than model choice does."]
A token is not a unit of work
Here is the part that makes agentic budgets behave strangely. An agent re-sends its accumulated context on every turn, so the tokens billed are not the tokens of the conversation; they are the tokens of every prefix of the conversation, summed.
Let the stable prefix be \(P\) tokens of system instructions and tool definitions, and let each turn append \(a\) tokens of tool result and prior assistant output. Turn \(t\) sends \(P + (t-1)a\) input tokens, so over \(n\) turns
The quadratic term is the whole story. Doubling the turn count roughly quadruples the input bill, so a task that goes badly at 20 turns instead of 5 costs about sixteen times as much in input tokens, not four. Nothing in the price list hints at this, because the list is linear in tokens and the tokens are quadratic in turns.
Now add incremental caching, the way a real harness uses it. On turn \(t\) everything except the newest \(a\) tokens is already cached and reads at \(r = 0.1\), while the new \(a\) tokens are written at \(w = 1.25\); turn 1 writes the prefix. The billed token-equivalents are
Compare the quadratic coefficients: \(a/2\) uncached against \(ra/2 = a/20\) cached. Caching divides the quadratic coefficient by exactly ten and does not touch the exponent. It is a large constant-factor win on a growth law it cannot change, which is why "we turned on prompt caching" is a real saving and never a fix for a loop that runs long.
With \(P = 8{,}000\), \(a = 1{,}600\) and \(n = 40\):
- Uncached: \(8{,}000 \times 40 + 1{,}600 \times 780 = 320{,}000 + 1{,}248{,}000 = 1{,}568{,}000\) input tokens. At $2 per million, $3.14.
- Cached: \(10{,}000 + 31{,}200 + 118{,}560 + 78{,}000 = 237{,}760\) token-equivalents. At $2 per million, $0.48.
A 6.6x reduction, and the cached 40-turn loop still bills 2.6 times what a cached 20-turn loop would. The lever with the better exponent is turn count.
[IMAGE: Line chart of cumulative billed input tokens against turn number, 1 to 40, with two curves: uncached rising to 1.568 million and incrementally cached rising to about 238,000. Both visibly curve upward rather than straightening. Caption: "Caching moves the curve down by a factor of 6.6 and leaves its shape untouched."]
flowchart TB
subgraph UC["Uncached loop"]
U1["Turn 1: 8k"] --> U2["Turn 20: 38.4k"]
U2 --> U3["Turn 40: 70.4k"]
U3 --> UT["Total 1,568,000 tokens"]
end
subgraph CA["Incrementally cached loop"]
C1["Turn 1: write 8k"] --> C2["Turn 20: read 36.8k, write 1.6k"]
C2 --> C3["Turn 40: read 68.8k, write 1.6k"]
C3 --> CT["Total 237,760 equivalents"]
end
UT --> R["Same O of n squared growth"]
CT --> R
classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
class U1,U2,U3,UT rose
class C1,C2,C3,CT emerald
class R slateA token is not the whole bill
Some of what you pay is not metered in tokens at all, and those terms grow with decisions rather than with text. Server-side search bills $10 per 1,000 searches however many results return. Code execution bills container time at $0.05 per hour beyond a monthly allowance of 1,550 hours, with a five-minute minimum per run, so two hundred three-second calls a day bill as roughly 16.7 hours rather than ten minutes. Managed agent sessions add $0.08 per session-hour, metered only while actually running. Pinning inference to one geography multiplies every token category by 1.1 (figures from Anthropic, Pricing, September 2026).
The subtlest term is still measured in tokens and still invisible to a per-request estimate: tool definitions are re-sent every turn. A browser toolset is documented at roughly 6,600 input tokens per request and a computer-use toolset at roughly 4,500, before any tool is called. Over a 40-turn loop that is 264,000 billed input tokens, $0.53 at $2 per million, for text the agent has already read 39 times. At the front of a cached prefix it costs about five cents; placed after something volatile, it is a line item nobody can find.
[IMAGE: Waterfall chart of one agent task's cost, starting at zero and adding bars for cached input, fresh input, output tokens, tool definitions, search fees and container time, with the non-token bars shaded amber. Caption: "Four of the six bars are invisible to a dollars-per-million-tokens estimate."]
A token is not a choice you control
Finally, the quantity of tokens a request consumes is partly the model's decision. Reasoning tokens bill at output rates, and their volume is set by the model's own behaviour under a requested effort level. Benchmarking 53 models across 14 basic arithmetic tasks, reasoning variants emitted roughly 18 times more tokens, sometimes at lower accuracy, with accuracy collapsing by up to about 36 percent when tokens were constrained and with non-monotonic accuracy-verbosity curves throughout (Srivastava et al., 2025, Do LLMs Overthink Basic Math Reasoning?, arXiv:2507.04023). Turning the effort dial up is therefore not reliably a purchase of accuracy: the harness behind the Holistic Agent Leaderboard found that higher reasoning effort reduced accuracy in the majority of its runs (Kapoor et al., 2026, Holistic Agent Leaderboard, ICLR, arXiv:2510.11977).
Fan-out is the same story at a larger grain. Anthropic measured agents using about four times the tokens of chat interactions and multi-agent systems about fifteen times, and found that token usage alone explained 80 percent of the performance variance on a browsing evaluation (Anthropic, 2025, How we built our multi-agent research system). Read that carefully: the variable driving quality is the variable driving spend. A token budget is a quality budget, and pretending otherwise is how cost programmes quietly ship regressions.
Seeing It in Motion
One turn of an agent loop touches four meters. Watching them tick separately is the fastest way to internalise why a single blended rate cannot describe the request.
sequenceDiagram
participant A as Agent harness
participant M as Model API
participant T as Server tool
A->>M: Send prefix plus history
Note over M: Cached portion reads at 0.1x<br/>New delta writes at 1.25x
M-->>A: Assistant turn plus tool call
Note over M: Output and thinking bill at output rate
A->>T: Execute search
T-->>A: Results plus per-call fee
Note over A,T: 10 dollars per 1,000 searches
A->>M: Resend everything plus results
Note over M: Context grew; next turn costs moreFour meters, three rates, one request. The usage object in each response is the only place all four appear together, which is why recording it per request is the difference between a cost model and a guess.
[IMAGE: Annotated screenshot-style mock of a single API response usage object, with callout arrows labelling input_tokens, cache_read_input_tokens, cache_creation_input_tokens, output_tokens and server_tool_use. Caption: "The five fields a cost model needs, and the one place they are reported together."]
Watch It Run
By the Numbers
| Quantity | Value | Where it comes from |
|---|---|---|
| Cache read multiplier | 0.1x of base input, 0.05x on some models | Anthropic pricing, Sept 2026 |
| Cache write multipliers | 1.25x (5 min), 2.0x (1 hour) | Anthropic pricing |
| Asynchronous batch discount | 50% on input and output | Anthropic pricing; matched by other major APIs |
| Pinned data residency | 1.1x on every token category | Anthropic pricing |
| Server-side search | $10 per 1,000 searches | Anthropic pricing |
| Code execution | $0.05 per container-hour, 5-minute minimum, 1,550 free hours monthly | Anthropic pricing |
| Managed agent session runtime | $0.08 per session-hour, running time only | Anthropic pricing |
| Browser toolset definition overhead | about 6,600 input tokens per request | Anthropic pricing |
| Tokeniser change within one family | about 30% more tokens for the same text | Anthropic pricing |
| Agent vs chat token usage | about 4x; multi-agent about 15x | Anthropic engineering blog |
| Share of BrowseComp variance from token usage | 80% | Anthropic engineering blog |
| Reasoning token multiplier | about 18x, sometimes at lower accuracy | Srivastava et al., arXiv:2507.04023 |
| Annual price decline per capability level | 9x to 900x | Epoch AI |
| Google monthly token throughput | 9.7T (2024) to 3.2 quadrillion (2026) | Alphabet, I/O 2026 |
| Enterprise AI market | $37B in 2025, 3.2x year on year | Menlo Ventures |
| Cost of validating one agent eval harness | about $40,000 for 21,730 rollouts | Kapoor et al., arXiv:2510.11977 |
Sources: multipliers and SKUs from Anthropic's published pricing as consulted in September 2026, which moves and should be re-checked rather than recalled. Token-usage multipliers from Anthropic's multi-agent write-up; reasoning figures from arXiv:2507.04023; price declines from Epoch AI; market size from Menlo Ventures; evaluation cost from arXiv:2510.11977. All vendor figures are list prices, not what a negotiated contract pays.
A Concrete Example
A support organisation handles 1 million tickets a month. A human touch costs $3.50 fully loaded. Three architectures are on the table, and the interesting result is not which one wins but by how much less than expected.
Shared assumptions: a 6,000-token prefix of system instructions, policy text and tool definitions; each agent turn appends 1,200 tokens of tool result and prior output and emits 350 tokens; two server-side searches per ticket at $10 per 1,000. Prices are the September 2026 list rates used above.
Step 1. Architecture A, one classification call on a small model. Input is the 6,000-token prefix plus a 900-token ticket; output 250 tokens. At $1 and $5 per million:
Under a cent per ticket, and it resolves 38 percent of them without a person.
Step 2. Architecture B, an eight-turn agent, uncached, on a mid-tier model. Input tokens follow the loop law: \(6{,}000 \times 8 + 1{,}200 \times 28 = 48{,}000 + 33{,}600 = 81{,}600\). Output is \(8 \times 350 = 2{,}800\). At $2 and $10 per million, plus $0.02 of search fees:
Twenty-six times the per-call cost of Architecture A. It resolves 71 percent.
Step 3. Architecture C, the same agent with an incrementally cached prefix. Using \(T_{\text{eff}}\) with \(P = 6{,}000\), \(a = 1{,}200\), \(n = 8\): \(7{,}500 + 4{,}200 + 2{,}520 + 10{,}500 = 24{,}720\) token-equivalents. So
The token bill fell by 2.2 times. Hold that number.
Step 4. Put the escalations back in. The cost that matters is per resolved ticket including the humans who handle the residue:
| Architecture | Model cost per ticket | Success \(s\) | Human cost per ticket | Blended | Monthly at 1M tickets |
|---|---|---|---|---|---|
| A, single call | $0.0082 | 38% | $2.170 | $2.178 | $2.18M |
| B, 8-turn agent, uncached | $0.2112 | 71% | $1.015 | $1.226 | $1.23M |
| C, same agent, cached | $0.0974 | 71% | $1.015 | $1.112 | $1.11M |
| C at 80% success | $0.0974 | 80% | $0.700 | $0.797 | $0.80M |
Step 5. Read the table properly. Architecture A is 26 times cheaper per call and the most expensive option by a wide margin. Caching, which cut the token bill by 2.2x, cut the blended cost by 9 percent. Nine percentage points of extra success rate saved $0.315 per ticket, more than three times the entire caching win. At this operating point the token bill is not the lever, and a cost programme that optimises it is optimising a term worth a tenth of the one next to it.
That conclusion flips when escalation is cheap or absent. If the residue costs $0.20 rather than $3.50, the blended figures become $0.132, $0.269 and $0.155, Architecture A wins, and the token bill is suddenly most of the cost. The ranking of architectures is a function of the cost of being wrong, which is the formal result in Zellinger and Thomson, 2025, Economic Evaluation of LLMs, arXiv:2507.03834: reasoning models start paying once a mistake costs more than roughly a cent, and a single large model overtakes a cascade once a mistake costs about ten cents. Until you price an error, "cheaper" is not a well-formed claim.
[IMAGE: Grouped bar chart of blended cost per resolved ticket for architectures A, B, C and C-at-80-percent, each bar split into model cost and human escalation cost. A second small panel repeats the chart with escalation priced at 20 cents, showing the ranking invert. Caption: "The same three architectures, ranked twice, by the cost of being wrong."]
Where It Breaks
The effective rate is a moving average of your own behaviour
\(p_{\text{eff}}\) is not a price you were quoted; it is an outcome of prompt structure, traffic shape and hit rate, so it moves when nobody changed a contract. A refactor that inserts a per-request identifier above a stable policy block can take a workload from 0.26 of list to 1.25, a 4.8x increase, and it looks like a traffic anomaly on every dashboard because token counts barely move. Cache hit rate belongs next to latency and error rate in the service's own metrics, not in a monthly finance review.
A price list cannot be diffed
Two vendors publishing the same dollars per million tokens are not selling the same thing if their tokenisers, cache multipliers and free allowances differ, and one bills server-side tools per call while the other folds that capability into tokens. No arithmetic makes those comparable on paper. The only comparable quantity is measured cost per completed task on your own traffic, which is why a migration assessment is an engineering project rather than a spreadsheet. See cost per completed task.
Mean cost per request is the wrong forecasting input
Agentic cost distributions are heavy-tailed, because turn count enters quadratically, tool result sizes are chosen outside the system, and retries restart loops. If one percent of runs cost a hundred times the median and the rest cost the median, spend per run is \(0.99 + 1.00 = 1.99\) median-units and that one percent is 50.3 percent of the bill. The mean is then a tail artefact: it moves when rare runs get slightly longer. Forecast p50, p90 and p99 separately and treat p99/p50 as a leading indicator, since it moves before quality metrics do. See cost variance and the tail of agentic spend.
The unit changes under you
A tokeniser change, a new default reasoning effort, a deprecation that forces a migration, a toolset version whose definitions grew: each reprices existing traffic without any price changing. The defence is a pinned, versioned cost benchmark over a fixed set of representative requests, reporting tokens and dollars per completed task for the current model and harness. Without it you cannot separate a vendor-side change from your own regression, and both get attributed to whichever team shipped most recently.
Non-token meters resist attribution
A per-request token count can be attributed to a feature, a customer and a tenant. A container-hour with a five-minute minimum cannot be, cleanly: it was consumed by whichever run happened to fall inside the window. Free allowances make it worse by setting the apparent marginal price to zero until the month you cross the threshold, at which point the fitted model is wrong by the full list rate. Treat an allowance as a threshold with a forecast crossing date. See token accounting and cost attribution and non-token charges in an AI bill.
Cost-blind evaluation selects the expensive design
If your evaluation harness records accuracy and not cost, it will reliably prefer the more expensive architecture, because more tokens usually buy some accuracy. This was measured directly: on HumanEval a simple baseline matched a published state-of-the-art agent architecture at roughly 2 percent of its cost, a difference invisible on an accuracy-only leaderboard (Kapoor et al., 2024, AI Agents That Matter, arXiv:2407.01502). Recording cost per run is cheap. Not recording it means every architectural decision has a hidden thumb on the scale.
Alternative Designs
| Design | How it works | Key advantage | Key limitation | Best when |
|---|---|---|---|---|
| List price per token | Expected tokens times the rate card | Trivial; vendor-published | Blind to cache mix, non-token SKUs, tokeniser drift and success | Single-turn, single-model, stable workloads |
| Effective rate per token | Weight the rate card by the fresh, written and read mix | Captures the largest structural discount | Still linear in a quantity quadratic in turns | Cache-heavy serving with stable prompts |
| Cost per completed task | Measure tokens, fees and success on a golden set, then divide | Comparable across vendors and architectures | Needs a trusted rubric and an evaluation budget | Any model or architecture choice |
| Cost per business outcome | Put escalation and rework in the denominator | The only figure a product owner can act on | Requires owning the downstream process | Features replacing a costed human workflow |
| Reserved capacity | Buy throughput, not tokens | Converts a variable bill into a fixed one | Idle capacity is pure loss | Steady base load at scale |
Two of these are complements rather than alternatives. An effective-rate model is how you drive one design's cost down; cost per completed task is how you choose between designs. Running the first without the second is how teams arrive at a beautifully cached agent that loses to a two-line classifier. Owning the hardware instead is a separate calculation, worked in self-hosting versus API break-even.
[IMAGE: Decision flowchart. Start: "Which cost unit should I use?" Branches on whether the workload is single-turn, whether escalation is costed, whether you are choosing between models, and whether utilisation is flat, terminating in the six rows of the table above. Caption: "The unit follows the decision, not the invoice."]
How It Is Used in Practice
Public leaderboards moved first, because they had the strongest incentive. The SWE-bench leaderboards now carry per-instance cost next to resolution rate, which changes what a top position means. The Holistic Agent Leaderboard went further and made cost-controlled evaluation the infrastructure rather than an annotation, orchestrating 21,730 rollouts across nine models and nine benchmarks for about $40,000 (Kapoor et al., 2026, arXiv:2510.11977). Artificial Analysis reports a cost-to-run figure for its index built from each model's input, cache, reasoning and answer token prices and the tokens actually used, which is \(C_{\text{task}}\) with the success term set aside.
Inside product teams the practice that sticks is duller: record the full usage object on every request, tag it by feature, tenant and purpose, and keep a pinned benchmark of dollars per completed task that reruns on every model, prompt or harness change. That is what lets a team say whether the bill moved because traffic changed, because the workload changed, or because the vendor changed something. Teams that store only aggregate spend can answer none of the three.
The environmental disclosures now emerging are the same problem in another currency. Google reports a median Gemini text prompt at 0.24 Wh, 0.03 gCO2e and 0.26 mL of water, and its own narrower boundary, counting active accelerator draw alone, gives 0.10 Wh for the same prompt (Google Cloud, 2025, Measuring the environmental impact of AI inference). A factor of 2.4 from the accounting boundary, before any question of model or grid. Per-request energy is as slippery a unit as per-request cost, for the same reasons. See per-request energy, water and carbon.
Insights Worth Remembering
-
The tokeniser is part of the price. A vendor can raise your bill by a third without changing a published number, and did. Cross-generation and cross-vendor comparisons have to measure token counts on your own text rather than inherit them.
-
Prompt caching is a constant-factor win on a quadratic law. It divides the quadratic coefficient by ten and leaves the exponent alone: the highest-leverage change available without touching behaviour, and never a substitute for shortening the loop.
-
Prefix stability is an interface with a price. One per-request value inserted above a stable block can multiply the effective input rate by nearly five. Treat prompt ordering as a contract and cache hit rate as a service metric.
-
Tool definitions bill per turn, not per use. A 6,600-token toolset over 40 turns costs more than most whole requests. Load capability on demand and keep whatever must be declared inside the cached prefix.
-
The variable driving quality is the variable driving spend. Token usage explained 80 percent of the performance variance on a browsing evaluation, so a token budget is a quality budget and a cap must be justified against completion rate.
-
The ranking of architectures is a function of the cost of being wrong. Price the error first. Reasoning models, cascades and single large models each win in a different regime of that one parameter, and token-level optimisation does not move you between regimes.
-
Optimise the largest term, not the most visible one. In the worked example a 2.2x cut in the token bill moved total cost by 9 percent, while nine points of success rate moved it by 28 percent. Token spend is the term with a dashboard, rarely the term with the money.
Open Questions
Can an outcome price be audited? Some vendors and applications now quote per resolved ticket rather than per token, which demonstrably moves tail risk from buyer to seller. Whether a buyer can verify an outcome claim without instrumenting the work themselves is unresolved, and the first generation of such contracts is likely to be priced from token models anyway, with margin covering the variance.
Does the quadratic loop law survive better context management? Compaction, external note-taking and code-execution tool interfaces all shrink what must be re-sent; one published workflow fell from 150,000 tokens to 2,000 by keeping results out of the window entirely. Whether these change the exponent or only the constants for realistic agents is open; the mechanism suggests they truncate the sum rather than flatten it.
What is the right unit for a heavy-tailed bill? Percentiles forecast well and price badly, because you cannot invoice a p99, and reserved capacity converts the problem into a utilisation problem rather than solving it. How a multi-tenant product should charge for a workload whose per-request cost spans two orders of magnitude is unsettled.
Will tokenisers converge or diverge? Fertility differences across languages are measured and consequential, and at least one proposal argues for fairer pricing rather than better modelling. Whether providers normalise against a language-independent unit or the divergence widens as tokenisers specialise is speculation, with commercial pressure pointing both ways.
Can cost-controlled evaluation hold as a norm? Recording cost per run is cheap and the evidence that omitting it distorts architecture choices is strong. Whether the practice survives the first time a cost column makes a favoured system look bad is sociological rather than technical, and the history of accuracy-only leaderboards is not encouraging.
Sources and Further Reading
- Anthropic. "Pricing." Claude Platform documentation, consulted September 2026. platform.claude.com/docs/en/about-claude/pricing
- Anthropic. "How we built our multi-agent research system." Anthropic Engineering, 2025. anthropic.com/engineering/multi-agent-research-system
- Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). "AI Agents That Matter." arXiv:2407.01502
- Kapoor, S., Stroebl, B., Kirgis, P., et al. (2026). "Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation." ICLR 2026. arXiv:2510.11977
- Srivastava, G., Hussain, A., Srinivasan, S., & Wang, X. (2025). "Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models." arXiv:2507.04023
- Zellinger, M. J., & Thomson, M. (2025). "Economic Evaluation of LLMs." arXiv:2507.03834
- Wani, S. G., Dholakia, A., & Ellison, D. (2026). "The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts." arXiv:2608.26235
- Lundin, J. M., et al. (2025). "The Token Tax: Systematic Bias in Multilingual Tokenization." arXiv:2509.05486
- Bergemann, D., Bonatti, A., & Smolin, A. (2025). "The Economics of Large Language Models: Token Allocation, Fine-Tuning, and Optimal Pricing." Proceedings of the 26th ACM Conference on Economics and Computation (EC '25). doi:10.1145/3736252.3742625
- Cottier, B., et al. (2025). "LLM inference prices have fallen rapidly but unequally across tasks." Epoch AI. epoch.ai/data-insights/llm-inference-price-trends
- Menlo Ventures (2025). "2025: The State of Generative AI in the Enterprise." menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise
- Google Cloud (2025). "Measuring the environmental impact of AI inference." cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference
- Pichai, S. (2026). "Google I/O 2026: opening keynote." Google blog. blog.google/innovation-and-ai/sundar-pichai-io-2026
- Pichai, S. (2026). "Alphabet earnings call Q2 2026: Sundar Pichai remarks." Google blog. blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2026
- Anthropic (2025). "Code execution with MCP: building more efficient agents." Anthropic Engineering. anthropic.com/engineering/code-execution-with-mcp
- SWE-bench leaderboards, which report per-instance cost alongside resolution rate. swebench.com
Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.