LLM Application Architecture intermediate 8 min read 7 flashcards

The Tool Layer as a Dependency Tier

How to budget timeouts across the tools inside a single model turn, why a tool failure belongs in the conversation rather than in an exception, and what tool definitions and results cost in context and cache.

An answer took 70 seconds. The model accounted for nine of them. The rest belonged to a ticket-system call that retried twice behind the scenes, a search index that returned 40 KB of JSON, and a tool that failed in a way the model read as an invitation to try something else. Tool calling is framed as a model capability. Operationally it is a fan-out to every flaky dependency your company owns, with a language model deciding the call graph at runtime.

One turn, many dependencies, one deadline

The defence against a slow dependency taking down everything above it is deadline propagation: set one absolute deadline at the edge and have each hop subtract the time already spent instead of applying a local constant. The Google SRE book treats this as a primary mitigation for cascading failure, and names missed RPC deadlines as the mechanism by which an overloaded server turns latency into wasted work and retries (Beyer et al., 2016, Site Reliability Engineering, ch. 22).

An agent turn makes the arithmetic unavoidable, because the model is in the loop several times. Suppose a 30-second end-to-end budget, three model hops at roughly 6 seconds each, and two tools in parallel per hop. The model consumes 18 seconds, leaving 12 for all tool work, so each hop's tools get 4 seconds, not the 30 or 60 that an HTTP client library uses if nobody sets it. Get this wrong in the generous direction and the first slow tool eats the budget, after which the model never produces an answer at all and the user sees a timeout instead of a degraded response.

Failure belongs in the conversation

A tool error is not an exception to raise. It is content to return, because the model decides what happens next. The Claude API gives it a shape: set is_error: true on the tool_result block and the model incorporates the failure into its response. The docs are specific about the content too, asking for what went wrong and what to try next, as in "Rate limit exceeded. Retry after 60 seconds.", rather than a generic "failed" (Anthropic, Handle tool calls).

Error text is therefore part of your prompt engineering, and retry placement becomes an architectural decision with a measurable cost. A retry inside the tool wrapper is invisible to the model and spends the deadline. A retry in the model loop is adaptive, because the model can switch tools, and expensive, because each attempt re-sends the conversation. The multiplier is real: an invalid or incomplete tool call is retried 2 to 3 times with corrections before the model gives up, so an ambiguous schema costs several turns of tokens per question. Declaring tools with strict: true removes that class of retry entirely by guaranteeing schema-valid inputs.

The worst failure shape is an empty success. A circuit breaker whose open state returns [] tells the model there is no data, and the model will say so confidently. An open breaker should return an error result saying the dependency is unavailable and not to retry this turn.

Definitions and results are a context budget

Everything here is paid for in tokens twice: once in the definitions sent on every request, once in the results that accumulate in the transcript.

Results dominate. Anthropic's worked example of moving tool orchestration into generated code takes a workflow from 150,000 tokens to 2,000, a 98.7% reduction, by letting the model write code against tool APIs so intermediate data never passes through the context window; a two-hour meeting transcript handed between tools is about 50,000 tokens on its own (Anthropic, 2025, Code execution with MCP). The response is to return high-signal fields and stable identifiers rather than whole objects, and to keep bulk data behind a handle (Anthropic, Define tools). Definitions are smaller but not free: schema input examples add roughly 20 to 50 tokens each, or 100 to 200 for nested objects, and a large tool surface is better served by deferred loading and tool search than by shipping every schema every time.

The tool layer is also part of the cache key, which makes schema changes a cost event. Definitions sit at the front of the prompt, so adding or reordering a tool invalidates the cached prefix for every conversation in flight, while changing tool_choice invalidates cached message blocks and leaves definitions and the system prompt cached. A deploy that adds one tool has a token bill attached.

When it breaks

Partial failure leaves an unanswerable turn. Every tool_use block needs a matching tool_result in the next message, so a worker that dies after two of three parallel calls leaves a conversation that cannot be continued until the missing results are synthesised as errors. Persist the pending set before executing, not after.

Side effects meet retries. Any tool that mutates the world needs an idempotency key derived from the call, not from the attempt; see orchestration, dependencies and idempotency and durable agent execution and recovery.

Errors in a 200 body. A tool that signals failure with {"error": "..."} and HTTP 200 is a success to your client and a failure to the model. Normalise at the wrapper so is_error and your metrics agree, or the dashboard will show a healthy tool the model has learned to avoid.

Shared tools are invisible single points of failure. One search tool behind every agent in the product has no alert of its own, and its degradation presents as "the assistant feels slow" across unrelated features. Give each tool its own latency and error SLO. Timeouts and breakers also say nothing about whether a tool's output should be believed; that is MCP server trust and tool poisoning.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Beyer et al., 2016, Site Reliability Engineering, ch. 22 sre.google
  2. Anthropic, Handle tool calls platform.claude.com
  3. Anthropic, 2025, Code execution with MCP anthropic.com
  4. Anthropic, Define tools platform.claude.com
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track