Non-Token Charges in an AI Bill
The line items a cost-per-token model cannot see, covering per-call tool fees, sandbox and session runtime, residency premiums and the token overhead of tool definitions, and why they decide the bill for agentic workloads.
An agent that runs thirty web searches while researching one question pays $0.30 in search fees at $10 per 1,000 searches (Anthropic, Pricing, figures as of September 2026). On a small model, that single line item can exceed everything the same task spends on tokens. A cost model built purely from input and output token counts will report a number that is confidently wrong, and it will be wrong by more as the workload becomes more agentic.
Unit economics of an AI feature assembles a cost per request. This concept enumerates the terms that are easiest to leave out of it.
The categories
Per-call server-side tool fees. Tools the provider executes are metered per invocation on top of tokens. Search is billed per query regardless of how many results come back; other server-side tools may be free of surcharge and still expensive in tokens, because the content they return enters the context. Web fetch carries no fee and can add 125,000 input tokens from one research PDF.
Sandbox and container time. Code execution is billed by execution time rather than by tokens, with a minimum billing increment and a monthly free allowance: a published example is 1,550 free container-hours a month and $0.05 per hour beyond it, with a five-minute minimum per run. A five-minute minimum is the part that bites. Two hundred short code-execution calls a day, each doing three seconds of real work, bill as roughly 16.7 hours.
Stateful session runtime. Managed agent sessions add a runtime charge, published at $0.08 per session-hour and metered only while the session is actually running, not while it sits idle waiting for input. This inverts the usual intuition: a long-lived agent is cheap when parked and expensive when thrashing, so a retry loop costs twice, in tokens and in wall-clock.
Routing and residency premiums. Pinning inference to one geography carries a 1.1x multiplier on every token category including cache reads, and regional endpoints on partner clouds carry a 10 percent premium over global ones. This is a compliance decision with a price, and it should appear as such in the business case rather than as an unexplained 10 percent variance.
The token cost of tool definitions. This is the one most cost models miss, because it looks like nothing. Declaring a toolset adds input tokens to every request in the loop: a browser toolset is documented at roughly 6,600 input tokens, a computer-use toolset at roughly 4,500, and the tool-use system prompt at several hundred more. Over a 40-turn loop, 6,600 tokens of definitions is 264,000 billed input tokens, or $0.53 at $2 per million, before the agent does anything. The fix is structural rather than frugal: put the definitions at the front of a cached prefix so they bill at 0.1x, and prefer loading capability on demand over declaring everything up front. Moving tool discovery out of the prompt entirely is the stronger version of the same move, and one published workflow fell from 150,000 tokens to 2,000 by keeping tool output out of the window rather than by compressing it (Anthropic, 2025, Code execution with MCP).
Why this changes the shape of the bill
Token charges are proportional to work. Most of the charges above are not. Per-call fees are proportional to decisions, minimum billing increments make short work expensive, and definition overhead is proportional to turns rather than to progress. A workload's blended cost per token therefore rises as it gets more agentic even when every price stays fixed, which is a large part of why forecasts built from a per-token rate drift upward against actuals.
The practical consequence is that the meter has to be the response, not the price list. Providers report tool invocations, cache reads and writes, and runtime in the usage object of each response; recording those fields per request is what makes the non-token terms attributable at all. See token accounting and cost attribution.
When it breaks
None of it ports. Per-call fees, runtime SKUs and free allowances differ by provider and have no common unit, so a multi-provider comparison built on token prices alone is not comparing the same product. Rebuild the cost model per provider on your own traffic.
Free allowances create cliffs, not discounts. A monthly allowance of container-hours makes the first portion of usage look free and the marginal unit look expensive at exactly the moment growth is fastest. Model the allowance as a threshold and forecast the month you cross it.
Fees survive failure. A search that returns nothing useful still bills, and a sandbox run that crashes still bills its minimum increment. Retry logic multiplies the non-token terms as reliably as it multiplies tokens.
Caching helps the definitions and not the fees. Moving a toolset into a cached prefix cuts its token cost by ten times; it does nothing for the per-call charges, which is why the two need separate controls: prompt structure for one, a decision budget for the other.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Anthropic, Pricing platform.claude.com
- Anthropic, 2025, Code execution with MCP anthropic.com
6 flashcards for this concept
Click a card to reveal the answer.