The Cost of Evaluation and Experimentation
Why evaluation is a standing line item that can rival production inference, the multipliers that make a golden set expensive, and how to buy statistical confidence at the lowest price.
Validating one agent evaluation harness cost about $40,000 in inference: 21,730 agent rollouts across nine models and nine benchmarks (Kapoor et al., 2026, Holistic Agent Leaderboard, ICLR, arXiv:2510.11977). That is a research budget, but the shape is the same inside a product team, and the shape is the problem: evaluation spend is not a project cost that ends at launch. It recurs on every model release, every prompt edit and every dependency bump, and unlike production inference it does not stop when traffic does.
Why it scales faster than production
The cost of one evaluation sweep is a product, not a sum:
Every factor grows for a good reason. Cases grow because each production incident adds a regression case. Repeats grow because a stochastic system needs several samples per case before a difference is distinguishable from noise; see error bars for LLM evals. Variants grow because you are comparing models, prompts and tool configurations at once. And \(C_{\text{run}}\) for an agentic task is not a single call but a full loop.
Put plausible numbers in: 300 cases, 5 repeats, 4 variants, $0.20 per agent run and $0.02 per judged case gives 6,000 runs at $1,320 per sweep. Run that on three occasions a week, which is modest for an actively developed feature, and it is roughly $206,000 a year. For a feature serving a few million requests a month at a cent each, evaluation is the same order of magnitude as serving.
Buying confidence efficiently
Four levers change the price of the same statistical conclusion, and all four are mechanical rather than clever.
The shared prefix across repeats is identical by construction, so cache reads at 0.1x of base input make repeats far cheaper than first runs; an eval harness that does not enable caching is paying full price for the most cacheable workload in the building. Offline sweeps have no latency requirement at all, which is exactly the trade the asynchronous batch tier exists for at 50 percent off both input and output; see the price menu. A two-stage design screens all variants on a cheap subset and promotes only survivors to the full set, which cuts the \(n_{\text{variants}}\) factor where it is most expensive. And paired designs, running every variant on the same cases with the same seeds and comparing within case, remove case difficulty as a variance source and reach significance with fewer repeats.
Reporting cost alongside accuracy is the discipline that makes all of this visible. Cost-controlled evaluation was proposed precisely because accuracy-only leaderboards reward expensive complexity: on HumanEval a simple baseline matched a published state-of-the-art agent at roughly 2 percent of its cost (Kapoor et al., 2024, AI Agents That Matter, arXiv:2407.01502). The same harness that spent $40,000 also found that raising reasoning effort reduced accuracy in the majority of runs, which is a result you can only obtain if the harness records what each run cost.
When it breaks
Cost pressure shrinks the golden set until the gate is noise. The cheapest way to cut an eval bill is to run fewer cases and fewer repeats, and that is also the fastest way to build a CI gate that blocks good changes and passes bad ones. If the budget forces a cut, cut variants or use a screening tier; do not cut repeats below what the observed variance requires. See LLM regression testing in CI.
The judge is a dependency with its own version. Changing the judge model, or its prompt, reprices and re-levels every historical result. Pin it, version it with the suite, and re-baseline deliberately rather than discovering the shift as an unexplained quality jump.
Eval spend hides inside production attribution. Running sweeps on the same API key as the feature makes evaluation look like traffic and inflates the feature's unit economics. Separate keys or cost tags per purpose are the minimum; see token accounting and cost attribution.
Agentic evals cost more than they look, because failures are the expensive runs. A task the agent cannot solve often burns the full turn budget before giving up, so the mean cost per case on a hard suite is dominated by failures. A suite that gets harder as you fix the easy cases gets more expensive per run over time even if nothing else changes.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.