An agent resolves support tickets in an average of 8 tool calls. A downstream inventory tool that used to answer in 80 ms starts taking 6 s but still returns correct results. Nothing errors and no circuit breaker trips. What happens, second by second?
Show the full answer Hide the answer
Second by second
Seconds 0 to 48. Eight tool calls at 6 s each is 48 s of waiting, plus model time, so a run that took about 12 s now takes about 60 s. No component reports a problem: the tool returns 200, the model returns valid output, the agent completes.
Around second 20. The user closes the tab. The HTTP connection drops. Unless cancellation is wired through the loop, the agent does not notice and keeps going - including the tool calls that create the ticket and send the confirmation email. The system is now performing side effects on behalf of a user who has left.
Around second 30. Concurrency, not throughput, becomes the constraint. Concurrent runs equal arrival rate times duration, so 20 runs per second at 60 s each is 1200 in flight where 240 were before. Thread pools, HTTP connection pools and any per-run memory footprint size for 240. The failure surfaces as connection-pool exhaustion in an unrelated service, which is where the investigation starts and where the cause is not.
Around second 45. The caller's own 30 s timeout fires and its retry begins the entire loop again. The first loop is still running. Two agents now work the same ticket, and because the tools are not idempotent the customer gets two emails.
Where it amplifies
The amplification is in the multiplication. A tool's latency is multiplied by the loop, and the loop length is decided by the model, so a p99 that is tolerable for a single request is intolerable inside an agent that may call the tool a dozen times on a hard ticket. A 6 s tool inside a 20-step research run is a two-minute request.
What the user sees
Nothing, for a long time. In a streamed interface the model emits nothing while a tool runs, so 6 s of silence per call reads as a hang, and the product has no way to distinguish "thinking" from "broken" unless tool starts and completions are surfaced as events.
What stops it, mechanically
- One deadline for the whole run, propagated into every tool call. Set it at the edge from the product's latency target, subtract elapsed time at each step, and refuse to start a tool call that cannot finish inside the remaining budget. This converts "slow" into a fast, visible failure.
- Cancellation that actually aborts. Client disconnect must cancel the loop's context, and the tool client must honour it, or abandoned work continues to consume the dependency.
- A per-tool latency budget with a degraded path. If the inventory tool exceeds 500 ms, the agent proceeds with "stock level unavailable" rather than waiting. A partial answer in 10 s beats a complete one in 60.
- Idempotency keys on every side-effecting tool, derived from the run id and step, so the duplicate loop's email is absorbed.
- Alert on age of oldest in-flight run and on tool p99 per tool, not on error rate, which never moves in this scenario.
What would have to be true for it to self-heal
Only if load shedding engages before pool exhaustion: an admission controller that caps concurrent runs and rejects new ones while the backlog drains. Without a bound on concurrency, the system has no mechanism that reduces offered load, so it degrades until something exhausts.
When not to spend the effort
A batch agent with no user waiting and idempotent tools can absorb a 75x latency change: the run takes longer and finishes. The controls above are bought by the presence of a waiting user or an irreversible side effect, not by the existence of an agent.