Cost per Completed Task
Why a list price per million tokens cannot rank models for a workload, how to compute cost per completed task instead, and the measured evidence that the cheaper model per token is often the more expensive one per outcome.
A model priced at $1 per million input tokens and $5 per million output tokens is half the price of one at $2 and $10. On a real agent workload the cheaper model can cost two and a half times as much per finished job, because it takes more turns to get there and finishes fewer of them. The per-token price is a true fact about the invoice and a poor predictor of the bill.
Unit economics of an AI feature builds a defensible cost per request. This concept is about the comparison that decides which model the request goes to, which needs a different denominator: not requests, but requests that succeeded.
The quantity to compute
For one workload and one candidate model, let \(T_{\text{in}}\) and \(T_{\text{out}}\) be the tokens the model actually consumes and emits to attempt a task, \(p_{\text{in}}\) and \(p_{\text{out}}\) the prices, \(F\) any per-call tool fees, and \(s\) the measured success rate against a rubric you trust:
Dividing by \(s\) is the step that reorders the candidates. It charges every successful outcome for the failed attempts that preceded it, which is what the business actually pays. Where a failure is escalated to a person rather than retried, the escalation cost belongs in the numerator as \((1-s) \cdot C_{\text{human}}\) instead, and the human term usually dominates everything else.
A worked comparison
Take an agent loop with an 8,000-token system-and-tools prefix, where each turn adds about 1,600 tokens of tool result and prior assistant output to the context and emits 400 tokens. Total input tokens billed over \(n\) turns is
because the whole context is re-sent every turn. Suppose the weaker model needs 14 turns and the stronger one 6.
- Weaker model, $1 / $5 per million: \(T_{\text{in}} = 257{,}600\), \(T_{\text{out}} = 5{,}600\), so an attempt costs $0.2576 + $0.028 = $0.286. At a 55 percent success rate, $0.52 per completed task.
- Stronger model, $2 / $10 per million: \(T_{\text{in}} = 72{,}000\), \(T_{\text{out}} = 2{,}400\), so an attempt costs $0.144 + $0.024 = $0.168. At 82 percent, $0.205 per completed task.
The cheap model is exactly half the price per token and 2.5 times the price per outcome. Nothing exotic produced that; the turn count enters quadratically and the success rate divides.
What the measurements show
This is not a thought experiment. Evaluations that control for cost keep finding that complexity bought accuracy at a price nobody had reported: on HumanEval a simple baseline matched a published state-of-the-art agent architecture at roughly 2 percent of its cost, which is invisible on an accuracy-only leaderboard (Kapoor et al., 2024, AI Agents That Matter, arXiv:2407.01502).
The reasoning dial has the same property. Benchmarking 53 models across 14 basic-arithmetic tasks found reasoning variants emitting roughly 18 times more tokens, sometimes at lower accuracy, with non-monotonic accuracy-verbosity curves (Srivastava et al., 2025, Do LLMs Overthink Basic Math Reasoning?, arXiv:2507.04023). A marginal formulation helps here: the Token Economy Score normalises a reasoning model's accuracy gain over a non-reasoning baseline by its generated-token multiplier, and across 151 model-benchmark runs found sequentially structured tasks give positive marginal efficiency while recall-heavy and already-saturated tasks do not (Wani et al., 2026, The Reasoning Tax, arXiv:2608.26235).
Where the threshold sits depends on what an error costs. Modelling the tradeoff as a single economic quantity, reasoning models become the better buy once the cost of a mistake exceeds roughly $0.01, and a single large model beats a cascade once a mistake costs about $0.10 (Zellinger and Thomson, 2025, Economic Evaluation of LLMs, arXiv:2507.03834).
When it breaks
The success rubric is the whole measurement. If \(s\) comes from a judge model, the judge's own errors and cost belong in the analysis, and a judge upgrade silently reprices every candidate. Ranking on a rubric nobody would accept in production is worse than ranking on price.
Token counts are not portable. \(T_{\text{in}}\) must be measured on your traffic with each model's own tokeniser. Vendors change tokenisers between generations, and a model family whose tokeniser emits roughly 30 percent more tokens for identical text is 30 percent more expensive at an unchanged headline price. See the tokenisation tax and token fertility across languages.
One number hides the distribution. \(C_{\text{task}}\) is a mean, and agentic cost distributions are heavy-tailed, so two models with the same mean can need very different budgets. See cost variance and the tail of agentic spend.
Measuring it is not free. A cost-per-task comparison across four models with enough repeats for usable error bars is itself a meaningful expense, which is why the cost of evaluation and experimentation is a line item rather than an afterthought.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Kapoor et al., 2024, AI Agents That Matter, arXiv:2407.01502 arxiv.org
- Srivastava et al., 2025, Do LLMs Overthink Basic Math Reasoning?, arXiv:2507.04023 arxiv.org
- Wani et al., 2026, The Reasoning Tax, arXiv:2608.26235 arxiv.org
- Zellinger and Thomson, 2025, Economic Evaluation of LLMs, arXiv:2507.03834 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.