Bounding what the agent may do
How production systems confine AI agents that read attacker-writable content while holding user credentials: the deterministic shell (OS sandbox, egress proxy, scoped tokens, human veto), the probabilistic scanners inside it, and the approval-binding failure class that defeated both across the 2025-26 incident record.
Eighteen months of CVE records, vendor security advisories, specification arguments and shipped enforcement code show how the industry actually bounds agent authority: a deterministic envelope the model cannot argue with, wrapped around best-effort scanners. The incidents (EchoLeak, the Amazon Q build-script compromise, MCPoison, two Claude Code sandbox advisories, mcp-remote, the GitHub MCP cross-repo exfiltration) sort into three failure classes, and the dominant one is not fooled humans but re-bound gates: approvals attached to file paths, bare hostnames, directory names and workspace settings that the attacker's side could rewrite. An architect leaves with a reference architecture, six decisions with their flip conditions, eight incident cards each ending in a transferable design rule, and a six-rung ladder for building and red-teaming the envelope.
In four of seven CVE-grade incidents the human-approval gate was never fooled and never even shown; it was re-bound, because the thing the gate checked (a path, a hostname, a directory name, a settings key) was writable by the agent or by repository content.
What you get out of it
- Every shipped defence converged on the same two-layer shape: deterministic enforcement outside the model (sandbox, egress proxy, scoped tokens, human veto) with probabilistic scanners strictly inside it; Google states the doctrine, Anthropic and OpenAI built it independently on the same OS primitives.
- The dominant failure class is gate re-binding, not approval fatigue: CVE-2025-54136 bound approval to a file path, GHSA-fg94 to a bare hostname with user-registrable paths, GHSA-7835 to a directory name the repo could define.
- Per-action approval fails at volume and the one measured number says how hard: capability envelopes removed 84% of permission prompts in Anthropic's internal usage, while the github-mcp-server reporter found the approval UI hiding actions behind 'See More'.
- The agent's token, not the user's request, defines the blast radius: the GitHub MCP exfiltration used only authorised calls, and the protocol's structural answer was the April 2025 rewrite making servers resource servers only, with token passthrough explicitly forbidden.
- No public postmortem yet documents a deterministic shell defeated through its front door; every recorded breach went around the boundary's definition, so the public record cannot tell you how the shells fail at design level, only that bindings are cheaper to attack.
Scope
Why this, now. Anthropic's June 2026 advisory pair showed the best-funded agent sandbox failing twice at the binding layer in one release, which turns the 2025 incident wave from anecdote into a named, recurring failure class worth designing against.
What it does not cover. Model-level jailbreak robustness, safety fine-tuning, multi-agent delegation, and every primary account outside this session's reachable network (Aim Security's EchoLeak write-up, Invariant Labs' GitHub MCP post, Willison's lethal-trifecta framing, Meta's Rule of Two, the Replit coverage, all arXiv PDFs and conference talks), which are named in the page as context and gaps rather than cited as evidence.
Other field guides
Keeping the training run alive
The only two complete public training logbooks (Meta's OPT-175B, BigScience's BLOOM) record what actually stops a big run: roughly two machine deaths…
23 sources · 12 organisations · 8 postmortemsEverything that demanded the whole program got archived: ten years of Meta's ML platform
A decade of one company's answer to a problem every deployment path has: how much of a dynamic program you capture ahead of time, and what you do wit…
24 sources · 11 organisations · 2 postmortemsKeep the format dumb: ten years of Hugging Face, measured from its own releases
Reads a decade of one company's decisions as a single argument about where a capability belongs, in the file format every reader must parse or in the…
20 sources · 9 organisations · 5 postmortems