Cost Variance and the Tail of Agentic Spend
Why the budget for an agent workload is set by the tail of its cost distribution rather than its mean, where the multipliers in that tail come from, and the controls that bound a single run without quietly lowering success.
Two features report the same mean cost per request. In the first, the most expensive one percent of requests cost about three times the median. In the second, they cost a hundred times the median. The second feature's mean is nearly twice its median, half of its entire spend comes from that one percent, and a month in which those runs get slightly longer blows the budget while every dashboard still reads normal.
That arithmetic is worth writing out, because it is the reason mean-based forecasts fail on agentic workloads. If 99 percent of runs cost one unit and 1 percent cost 100, total spend per run is \(0.99 + 1.00 = 1.99\) units, so the rare runs are 50.3 percent of the bill. The median tells you nothing about the bill, and the mean tells you nothing about the risk.
Where the tail comes from
Four mechanisms generate it, and they compose multiplicatively rather than adding.
Turn count enters quadratically. An agent re-sends its whole context on every turn, so total billed input tokens over \(n\) turns grow as \(O(n^2)\) even though the work per turn is constant. A task that needs 20 turns instead of 5 does not cost four times as much; it costs roughly sixteen times as much in input tokens. Prompt caching divides the quadratic term by ten but does not remove it.
Tool results are unbounded. Context arrives from outside the system, and its size is not the agent's choice. A vendor's own guidance puts a large documentation page at roughly 25,000 tokens and a research-paper PDF at roughly 125,000 (Anthropic, Pricing). One unlucky fetch can cost more than a hundred ordinary turns, which is why a maximum content size on every tool is a cost control and not a hygiene preference.
Fan-out multiplies. Measured against chat, agents use roughly four times the tokens and multi-agent systems roughly fifteen times, and token usage alone explained 80 percent of the performance variance on a browsing evaluation (Anthropic, 2025, How we built our multi-agent research system). The uncomfortable implication is that the thing driving quality is the thing driving spend, so a fan-out cap is a quality cap.
Reasoning effort is a per-request variable. Thinking tokens bill at output rates and their volume is decided by the model, not the caller, with non-monotonic returns (Srivastava et al., 2025, arXiv:2507.04023). Two identical requests can differ several-fold in output tokens for reasons no dashboard explains. See controlling reasoning length.
Budgeting a distribution
Report percentiles, not a mean: p50, p90, p99 and max cost per run, per feature and per customer. Forecast with the percentiles, because a headroom figure derived from a mean is wrong by whatever the tail ratio is. Track the tail ratio p99/p50 as a first-class metric; it moves when a prompt change alters loop behaviour, often before any quality metric moves.
Then bound the individual run. A per-session token budget, a maximum turn count, a maximum tool-result size and a wall-clock timeout together convert an unbounded distribution into a truncated one, which is the only form that can be capacity-planned. Fleet-level controls belong alongside them, not instead of them: see spend guardrails and quotas.
When it breaks
Truncation moves cost rather than removing it. Capping turns at the p90 kills the runs that needed 30 turns, and those were disproportionately the hard tasks. Cost per attempt falls, success falls, and cost per completed task can rise. A cap is defensible only when measured against completion, and the honest version degrades gracefully: hand off to a person, or return a partial result, rather than failing after spending the full budget.
Concentration is usually per-customer, not per-request. In a multi-tenant product one tenant's document sizes or query style can put them permanently in the tail. Attribution by customer is what surfaces this, and a per-tenant quota is the fix; see token accounting and cost attribution.
Retry policy is a tail generator. Exponential backoff with a generous retry count turns a transient tool failure into several full-context attempts. Retries that resume from saved state cost a fraction of retries that restart the loop.
The tail is where correlated failures live. Runaway loops, a tool that starts returning larger payloads, and a prompt regression that adds turns all show up first as a fattening tail and only later as an incident. Alert on the shape of the distribution, not on daily spend, which averages the signal away.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Anthropic, Pricing platform.claude.com
- Anthropic, 2025, How we built our multi-agent research system anthropic.com
- Srivastava et al., 2025, arXiv:2507.04023 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.