Subagent Context Isolation
also called Parallel Context Windows, Context Partitioning Across Agents
The property that makes multi-agent systems worth their cost - each subagent explores with its own context window and returns only a condensed result, so the system reads far more than one context could hold.
A research task needs 40 sources read. At 8000 tokens each that is 320000 tokens of reading for an answer that is two paragraphs long. One agent cannot hold it: the context fills, compaction starts discarding, and the answer is built from whatever survived.
Five subagents, each reading eight sources in its own context and returning a 500-token summary, let the lead agent reason over 2500 tokens that represent 320000. That compression is the actual product of a multi-agent design. It is not "more agents think better"; it is that separate context windows let the system read more than any single window holds.
Why it matters
Getting this wrong is expensive in a specific way. Teams adopt multi-agent architectures for tasks that fit comfortably in one context, and pay the coordination cost for nothing. The question "does this task exceed one context window, or need genuinely parallel reading" separates the designs that pay from the ones that are fashion.
The cost side is documented and large. Anthropic reports that its multi-agent research system used roughly 15 times the tokens of a chat interaction, with single-agent research at about 4 times. A design that multiplies spend by an order of magnitude needs a reason that survives being stated out loud.
Implementation patterns
- Give each subagent an objective, an output format, guidance on tools and sources, and explicit task boundaries. Anthropic's account names exactly these four, and the boundaries are what prevent overlap.
- Constrain the return payload. A subagent that returns its full transcript defeats the isolation: cap the result at a few hundred tokens in a fixed schema, with source identifiers so the lead can cite without re-reading.
- Partition the work by the lead agent, not by the subagents, since no subagent can see what another is doing.
- Keep writes out of the subagents. Workers read and propose; the lead executes side effects once, through deterministic code.
- Run them concurrently and bound the fan-out. Three to seven workers is the usual working range before coordination overhead and rate limits dominate.
Industry example
Anthropic's engineering write-up on building its multi-agent research system describes an orchestrator-worker pattern in which a lead agent coordinates and delegates to specialised subagents running in parallel, and reports that this configuration outperformed a single-agent baseline by 90.2% on its internal research evaluation. The same post is candid about the limits: domains requiring all agents to share the same context, or with many dependencies between agents, are a poor fit, and most coding tasks contain fewer genuinely parallelisable subtasks than research does.
Failure scenarios
- Overlap. Vague delegation sends three subagents at the same sources; the report gives the example of subagents duplicating work on a semiconductor-shortage query.
- Gaps. Between two boundaries nobody was told to cover, the answer is silently missing a source and reads as complete.
- Fluent wrong summaries. A worker that read the wrong document returns a confident condensation of it, and the lead has no way to tell - the isolation that saved tokens also removed the evidence.
- Context exhaustion in the lead. Returns that are not capped refill the very window the design existed to protect.
- Cost surprise. The 15x multiplier arrives on the invoice before anyone measures whether quality moved.
Trade-offs
Gained: reading capacity beyond one context, wall-clock reduction on genuinely parallel work, and per-worker specialisation of tools and prompts.
Paid: roughly an order of magnitude more tokens, a coordination layer with its own failure modes, harder debugging because no single transcript explains the answer, and a quality risk that is invisible - the lead cannot audit what a worker discarded.
When the bill arrives: at production volume, and during the first incident where an answer was wrong and nobody can reconstruct which worker went astray.
When not to use it
When the task fits in one context window, use one agent. When the subtasks depend on each other's intermediate findings, sequential steps in one context are both cheaper and better, because the dependency defeats parallelism anyway. When the work is a known sequence, code the sequence and use no agent at all. And when the task is coding: the same write-up notes that agents are not yet good at real-time delegation to each other, and that coding contains fewer parallelisable subtasks than research. The test to apply before adopting this: can you write down non-overlapping boundaries for each worker? If not, the task is not parallelisable and the fan-out will duplicate work at 15 times the price.
Interview question
Q: Your team proposes replacing a single support-triage agent with a supervisor and four specialised workers. What evidence would convince you, and what would you measure after?
What a strong answer covers: the deciding evidence is whether a triage case requires reading more than one context window can hold, and whether the four subtasks can be given disjoint boundaries. Neither is usually true for triage, so the likely answer is no. After any such change, measure tokens per resolved ticket, wall-clock p95, and resolution quality on a labelled set - and require the quality gain to justify the token multiplier, not merely to exist.
Quick check
Quiz: What does a subagent's separate context window buy that a longer single context does not? Parallel reading capacity - many windows' worth of source material compressed into short returns - rather than more attention over the same material.
Flashcard: What is the reported token cost of a multi-agent system relative to a chat interaction, and what must that cost buy? — About 15 times, per Anthropic's 2025 write-up; it has to buy reading capacity beyond one context or genuine parallelism, or it buys nothing.