Interview prompt. A supervisor agent fans out to five worker agents that each call tools with real side effects - creating tickets, sending emails, updating records. One worker fails at its fourth tool call, after two of those side effects have already happened. Tell me what the system does next.
Show the full answer Hide the answer
What the interviewer is testing
Whether you apply distributed-transaction thinking to agents, or treat the agent as a black box that either "worked" or "did not". A multi-agent run that performs side effects is a distributed workflow with a non-deterministic planner, and the interesting question is what state the world is in when the planner stops halfway.
The clarifying questions that change the answer
- Are the tools idempotent, and do they accept an idempotency key? If
create_ticketis safe to call twice with the same key, most of this problem disappears. If it is not, no retry strategy is safe. - Is there a durable run log? Not a transcript for debugging - a record per executed step with the step id, the tool, a hash of the arguments and the result, written before the caller sees the result.
- Which of the five workers' outputs does the final answer depend on? Partial success may be perfectly acceptable for a research fan-out and unacceptable for a fan-out that provisions accounts.
- Is a user waiting? A synchronous run has seconds to decide; an asynchronous one can park the failure for a human.
A strong answer's arc
- Resume, do not restart. Anthropic's write-up of its multi-agent research system makes the point directly: restarts are expensive and frustrating, so the system resumes from where the agent was when the error occurred, using checkpointing and retry logic around the model's own adaptability. Restarting a 12-step run also re-executes the two side effects that already succeeded.
- Replay through the run log with idempotency keys. Steps 1 to 3 are replayed from the log, not re-executed. Step 4 is retried with the same idempotency key, so a duplicate reaching the tool is absorbed.
- Classify tools by reversibility, and order the plan accordingly. Reversible and internal actions run early and unattended; irreversible external ones (sending the email, issuing the refund) run last, together, ideally behind one approval. This converts most partial failures into "nothing irreversible happened yet".
- Compensate what cannot be replayed. A sent email has no undo, so the compensation is a correction message, and the run log has to record enough to write one.
- Do not let the supervisor re-plan on stale context. After recovery it needs the updated world state, or it will plan around side effects that already exist.
Common weak answers
- "Retry the whole task." Duplicates side effects and is expensive: the same write-up reports multi-agent systems consuming roughly 15 times the tokens of a chat interaction, so a blind re-run of a fan-out is a real cost, not just a correctness bug.
- "Wrap it in a transaction." There is no transaction spanning a ticketing system, an email provider and a database. The saga is the only available shape.
- "Add a try/except and tell the model to handle it." The model is the component that failed to follow the plan; making it the recovery mechanism puts the unreliable part in charge of correctness.
What a strong answer adds
The observation that the safest multi-agent design is the one where workers cannot perform side effects at all. Workers read, search and propose; the supervisor executes the writes, once, through an ordinary deterministic code path with idempotency and an audit record. That gives up some parallelism on the write phase and removes the entire class of partial-failure reasoning above.
When this is over-engineering
A read-only research fan-out needs none of it: restart is cheaper than a saga when the only cost of a failed run is tokens. The machinery becomes mandatory the moment a worker can change something outside the process.