intermediate 3 min answer

You want an LLM-as-judge suite to gate every pull request that touches a prompt. It holds 400 cases; each case runs the product (about 3000 input and 400 output tokens) then a judge call (about 3500 input and 200 output). Roughly what does one run cost and how long does it take, and does the answer change the gate design?

llm-evaluationci-gatecostrate-limitsbatch
Show the full answer Hide the answer

The assumptions, stated

Per case: 6,500 input tokens and 600 output tokens across the two calls. For 400 cases that is 2.6M input and 0.24M output tokens per run. Prices below are Anthropic's published list rates in 2026 and are the kind of number that moves, so carry them as a variable, not a constant: a mid-tier model at $2 per million input and $10 per million output, a frontier model at $4 and $20.

The arithmetic

  • Mid-tier end to end: 2.6 × $2 = $5.20, plus 0.24 × $10 = $2.40 → about $8 a run.
  • Frontier end to end: 2.6 × $4 = $10.40, plus 0.24 × $20 = $4.80 → about $15 a run.
  • Mixed, which is what most teams run — product on the frontier model, judge on the mid-tier — lands between, near $12.

So a full run is single-digit to mid-teens dollars. At 10 prompt-touching runs a day that is roughly $80 to \(150 a day, **\)2k to $5k a month**, before reruns on flaky cases.

The number that actually constrains the design

Wall clock is not set by one call. A product call emitting 400 tokens takes a few seconds; the judge, less. Serially, 400 cases at roughly 8 seconds each is about 53 minutes — unusable as a gate. Raise concurrency to 20 and the arithmetic says three minutes, and that is where the real limit appears: 2.84M tokens in three minutes is close to a million tokens per minute, which is well above an ordinary account's throughput allowance. If the account's limit is 400k input tokens per minute, the run cannot finish in less than 2.6M ÷ 400k ≈ 6.5 minutes no matter how much concurrency you add, and every one of those tokens is taken from the same bucket production is drawing on.

Which assumption dominates the error

Output tokens, because they are priced several times input and because real answers overrun estimates. Second is whether the eval calls the real retriever: if it does, add its latency and its own cost, and the input-token figure becomes a distribution rather than a constant.

What the numbers rule in and out

  • Ruled out: the full 400-case suite on every push, sharing production's rate-limit bucket.
  • Ruled in: a tiered gate. A 40-case smoke suite per pull request costs under a dollar and finishes inside a minute. The full suite runs nightly and before release.
  • Ruled in: the batch path for the nightly run. Anthropic publishes batch processing at 50% of standard rates (2026), asynchronous, which suits a suite nobody is waiting on and roughly halves the largest line.
  • Ruled in: caching the stable prefix. Most of those 3,000 input tokens are the same system prompt on every case; cache it and the dominant cost line shrinks substantially.
  • Required either way: a separate key or project for eval traffic, so a hot evening of pull requests cannot rate-limit customers.

When this is the wrong answer

For a suite of 30 cases on a weekly release, all of this is over-built: run it serially, pay the $0.60, and spend the saved effort on better cases. The gate design only becomes interesting when the run cost or the run time exceeds what a developer will wait for, which is roughly ten minutes and roughly a dollar.