An agent resolves support tickets in up to 20 steps. The pod running one agent is evicted at step 12, after the agent has already created a ticket and sent a customer email. Walk through what happens next, and what must be true for the retry to be safe.
Show the full answer Hide the answer
Second by second
The container receives its termination signal and the process dies inside the grace period. The streaming response to the caller ends mid-token, so the client sees a truncated reply or a 502. The transcript — the goal, the twelve observations, and the record of which tools already ran — was a local variable, so it is gone. The orchestrator knows only that a run id stopped reporting.
Whatever retries next starts from step 1 with the same input. It re-reads the ticket, reaches the same conclusion, creates a second ticket, and sends a second email. The customer now has two messages that contradict each other on wording, and the queue has a duplicate.
Where it amplifies
The duplicate is now input. The next run reads the ticket system and finds two open tickets for one complaint, which looks like an escalation pattern, so a rule routes it to a human with higher priority — or a second agent picks up the duplicate and works it in parallel. Re-running also means a fresh step budget and a fresh token budget: the cost of the ticket is now doubled, and the run that was 60% finished contributed nothing.
What the user sees
Two emails, minutes apart, from the same assistant. Nothing in the logs says "we sent this twice", because each run believed it was the first.
What stops it
- The transcript is durable state, written after each step, keyed by run id. The agent loop is a workflow, not a function call: observation, chosen action and result are appended before the next step begins. Recovery replays the journal instead of the work. The cost is an extra write per step, commonly tens of kilobytes, and a few milliseconds of latency.
- Every side-effecting tool takes an idempotency key derived from
(run_id, step_index, tool_name, argument_hash), and the tool enforces it, not the agent. The model cannot be relied on to remember that it already sent the email; the email service can be relied on to reject a duplicate key. - Intent is journalled before execution and outcome after. That gives recovery three states rather than two: not started, done, and unknown. Prefer sending the unknown case to a human rather than retrying it, because a retry on "maybe sent" is how one failure becomes two customer-visible messages.
- Budgets survive the restart. Steps, tokens and wall clock are attributes of the run, not of the process, or a crash loop silently multiplies spend.
What would have to be true for it to self-heal
Every tool idempotent, the transcript durable, and the budget persistent. With all three, a restart costs about 2 seconds of replay and nothing else, and the failure never reaches the customer. Agents running in production without all three do not self-heal; they duplicate. With any one missing, recovery needs a person, and the honest design decision is to make the ambiguous case loud rather than automatic.
When this is the wrong answer
A read-only agent that searches and summarises needs none of this. Replaying it costs tokens and annoys the user, and durable per-step state would add latency to every step for no safety gain. The rule is sharp: durability becomes mandatory the moment a tool has an effect outside the process, and is optional until then. Teams get this backwards by building durable execution for their research agent and leaving the refund agent in memory.