pattern

Tool Result Budget

also called Tool Output Capping, Result Truncation Contract

A hard cap on the tokens any single tool may return into the model's context, with pagination and summarisation behind it, so that one unlucky query cannot fill the window and end the run.

tool-callingagentscontext-windowcostpagination

An agent calls search_orders(customer_id) on a customer with 4000 orders. The tool does what it was written to do and returns them all: 180000 tokens of JSON into a 200000-token context. The run ends there - either the request errors on context length, or the framework truncates the middle and the agent continues reasoning over a mutilated result it believes is complete.

The tool has no bug. It was designed for a caller that reads a return value into a variable, and it is now being called by a component whose input is metered and finite.

Why it matters

Tool results are the least controlled input to an agent's context. The system prompt is written once, the user turn is short, retrieved chunks are capped by design - and tool output is whatever a downstream system felt like returning, from any tool, at any step. In a loop, that output accumulates: every prior result is re-sent with every subsequent model call, so one oversized result is paid for on every remaining step of the run.

The cost is quadratic in the wrong place. A 30-step loop that admits a 20000-token result at step 5 re-sends it 25 times.

Implementation patterns

  • Cap at the tool boundary, in tokens, not rows. Count with the model's own tokeniser - Hugging Face's tokenizers library covers most open models and providers publish counting endpoints for their own - because rows, bytes and characters all mispredict tokens for JSON.
  • Return a page plus a cursor, and say so. {"results": [...], "returned": 20, "total": 4137, "cursor": "..."}. The model can then ask for more, and it knows it has not seen everything.
  • Budget per tool and per run. A per-run ceiling of, say, 40% of the context for all tool output keeps room for instructions, history and the answer.
  • Summarise server-side for large results rather than truncating: a cheap model condensing 4000 orders into a paragraph of statistics is more useful to the agent than the first 20 orders.
  • Make truncation explicit in the payload - a truncated: true field - so the model does not reason as though it saw everything.
  • Prefer aggregation tools to listing tools. count_orders and order_summary exist precisely so the agent does not need the list.

Industry example

The pattern is a direct descendant of pagination in public web APIs, and it exists for the same reason: an unbounded list endpoint eventually meets a caller with a large account. Tool-calling interfaces re-learned it the hard way because early implementations exposed existing internal functions directly, and internal functions written for another service rarely bound their output. Providers' own tool-use guidance now converges on the same advice - return concise structured results, paginate, and tell the model when there is more.

Failure scenarios

  • Silent truncation by the framework, which drops the middle of the conversation and takes early instructions with it.
  • Confident wrong answers from partial data. The agent reports "the customer has 20 orders" because that is what it received.
  • Loop death spiral. With the context nearly full, each step has less room, the model's answers get shorter and worse, and the loop runs out of room before it runs out of steps.
  • Cost blowout. One oversized result multiplied across the remaining steps is the single most common cause of an agent run costing an unexpected amount.
  • Prompt-cache invalidation. A large variable result early in the transcript means nothing after it can be cache-matched across turns.

Trade-offs

Capping buys a bounded, predictable context and a run cost that does not depend on which customer was queried. It pays in tool complexity - every listing tool now needs pagination, a token counter and a summary mode - and in extra round trips, since an agent that needs page two spends another model call to ask for it.

When the bill for not doing it arrives: in production, on the largest customer, which is the account you least want to give a wrong answer to.

When not to use it

When the tool's output is bounded by construction - a lookup returning one record, a boolean check, a status - a budget is ceremony. When the agent is a single-step call rather than a loop, the accumulation argument disappears and a generous cap is enough. And when the full result genuinely must be reasoned over, do not truncate it into the context: write it to a file or a table and give the agent a query tool over it, so the data stays outside the window and the agent pulls what it needs. The decision rule: cap whenever a tool's output size is controlled by data rather than by schema.

Interview question

Q: An agent's average run costs $0.12 and the p99 is $9. Where would you look first, and what would you change?

What a strong answer covers: the distribution shape says a small number of runs admit far more tokens than the rest, and the usual cause is one tool returning a data-dependent payload; the diagnosis is per-tool result-size percentiles logged per call, not per-run totals; the fix is a token cap with pagination on the offending tool plus a per-run tool-output budget; and the check afterwards is that p99 cost falls without a rise in task failure rate, which would indicate the cap is now too tight.

Quick check

Quiz: Why does a single 20000-token tool result at step 5 of a 30-step agent loop cost far more than 20000 tokens? Because the transcript is re-sent on every subsequent model call, so it is charged roughly 25 more times as input.

Flashcard: What should a listing tool return to a model instead of all 4137 matching rows? — A capped page with a cursor and the true total, plus an explicit truncated flag, so the model knows what it has not seen and can ask for more.