Evidence ledger 29 sources Checked 02 Oct 2026

Evidence ledger

One row per claim in Bounding what the agent may do: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

One row per claim. Tier grades follow the skill's vocabulary (postmortem, source, adr, casestudy, blog, paper, talk, vendor). Every URL below was fetched from this session on the check date. This session's network reaches a limited set of hosts (GitHub, GitLab, anthropic.com, claude.com docs, storage.googleapis.com, npm/PyPI registries); several primary accounts that exist outside that set are named in the guide as gaps rather than cited.

# Org Title Tier Published Checked URL Claim I take from it Supporting quote or figure
1 Microsoft (CNA) CVE-2025-32711 record, CVE List V5 postmortem 2025-06-11 2026-10-02 https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/32xxx/CVE-2025-32711.json M365 Copilot had a zero-interaction information-disclosure flaw driven by injected instructions ("EchoLeak"), CVSS 9.3 "Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network." baseScore 9.3, CRITICAL
2 AWS aws-toolkit-vscode commit 678851b postmortem 2025-07-13 2026-10-02 https://github.com/aws/aws-toolkit-vscode/commit/678851b The Amazon Q wiper incident entered through the build script: a commit titled as an inline-completion fix added a prod-only step that downloads a file from a stability ref and installs it as src/extensionNode.ts Commit title "fix(amazonq): should pass nextToken to Flare for Edits on acceptance…" vs. diff: scripts/package.ts +67 lines, downloads raw.githubusercontent.com/aws/aws-toolkit-vscode/stability/scripts/extensionNode.bk → src/extensionNode.ts, gated on STAGE === 'prod' and directory containing "amazonq"
3 Anthropic GHSA-fg94-h982-f3mm (CVE-2026-54316) postmortem 2026-06-13 2026-10-02 https://github.com/anthropics/claude-code/security/advisories/GHSA-fg94-h982-f3mm A pre-approved allowlist hostname that serves user-controlled content is an exfiltration channel; CVSS 6.0; patched 2.1.163 "Because the hostname huggingface.co was pre-approved as a bare hostname for the WebFetch tool, any path on that domain—including attacker-controlled model repositories—was auto-approved without a permission prompt … creating a covert out-of-band channel for encoding and exfiltrating data" (CWE-183, CWE-515)
4 Anthropic GHSA-7835-87q9-rgvv (CVE-2026-55607) postmortem 2026-06-25 2026-10-02 https://github.com/anthropics/claude-code/security/advisories/GHSA-7835-87q9-rgvv The sandbox trusted a path identity the repository could redefine: worktrees named .git plus symlinks let writes land outside the sandbox (~/.zshenv), CVSS 7.7; patched 2.1.163 "worktree handling allowed creation of worktrees named '.git' … an attacker could overwrite files in the user's home directory (such as .zshenv), leading to code execution outside of seatbelt sandbox restrictions." / "Reliably exploiting this required the user to clone a malicious repository containing prompt injection content and run Claude Code against it."
5 JFrog (CNA) CVE-2025-6514 record, CVE List V5 postmortem 2025-07-09 2026-10-02 https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/6xxx/CVE-2025-6514.json The client-side OAuth helper executed attacker input from the server it was authenticating to, CVSS 9.6 "mcp-remote is exposed to OS command injection when connecting to untrusted MCP servers due to crafted input from the authorization_endpoint response URL" baseScore 9.6, CRITICAL
6 npm registry mcp-remote version timeline source 2025-06-17 2026-10-02 https://registry.npmjs.org/mcp-remote The fixed release 0.1.16 shipped 2025-06-17, three weeks before the CVE record published registry time map: "0.1.16": 2025-06-17 (CVE-2025-6514 record published 2025-07-09)
7 Cursor (via GitHub CNA) CVE-2025-54136 record, CVE List V5 postmortem 2025-08-01 2026-10-02 https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/54xxx/CVE-2025-54136.json Cursor bound MCP approval to the config file's existence, not its content: editing an already-approved mcp.json re-ran silently, CVSS 7.2; fixed in 1.3 "Once a collaborator accepts a harmless MCP, the attacker can silently swap it for a malicious command (e.g., calc.exe) without triggering any warning or re-prompt."
8 Microsoft (CNA) CVE-2025-53773 record, CVE List V5 postmortem 2025-08-12 2026-10-02 https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/53xxx/CVE-2025-53773.json GitHub Copilot / Visual Studio shipped a prompt-injection-to-local-code-execution flaw, CVSS 7.8 "Improper neutralization of special elements used in a command ('command injection') in GitHub Copilot and Visual Studio allows an unauthorized attacker to execute code locally." baseScore 7.8, HIGH
9 Anthropic (via GitHub CNA) CVE-2025-52882 record, CVE List V5 postmortem 2025-06-24 2026-10-02 https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/52xxx/CVE-2025-52882.json The IDE companion opened a control channel any webpage could reach: unauthorized websocket connections to the Claude Code extension, CVSS 8.8 "Claude Code extensions in VSCode and forks … are vulnerable to unauthorized websocket connections from an attacker when visiting attacker-controlled webpages." patched June 13th, 2025
10 GitHub community github-mcp-server issue #844 postmortem 2025-08-08 2026-10-02 https://github.com/github/github-mcp-server/issues/844 The cross-repo exfiltration flow (malicious public issue → agent reads private repos → agent publishes to public repo) was reproduced and reported against the official server; the issue closed stale with no documented fix Reporter (sei-renae) describes prompt injection via public issues leaking private-repo data; notes approval UI showed a prominent "Continue" while actions sat behind "See More"; references the Invariant Labs May 26, 2025 publication; closed as stale
11 GitHub github-mcp-server README source 2026 (checked) 2026-10-02 https://github.com/github/github-mcp-server The official server's mitigations are capability cuts: toolset allowlists, --read-only with priority over tool requests "Read-only mode takes priority: write tools are skipped if --read-only is set, even if explicitly requested via --tools" / "Enabling only the toolsets that you need can help the LLM with tool choice and reduce the context size."
12 MCP project PR #284 "[RFC] Update the Authorization specification for MCP servers" adr 2025-04-07 2026-10-02 https://github.com/modelcontextprotocol/modelcontextprotocol/pull/284 The community rewrote MCP auth away from every-server-an-OAuth-provider after a recorded argument (114 comments); closed unmerged in favour of #338 "Malicious actors could deploy fake MCP Servers mimicking our MCP Server. If victims are tricked into configuring these rogue servers, attackers could obtain valid AccessTokens."
13 MCP project PR #338 (merged authorization split) adr 2025-04-23 2026-10-02 https://github.com/modelcontextprotocol/modelcontextprotocol/pull/338 The merged fix makes MCP servers resource servers only, with RFC 9728 metadata discovery Merged 2025-04-23; "MCP Servers are Resource Servers only"; adopts OAuth 2.0 Protected Resource Metadata (RFC 9728)
14 MCP project Security Best Practices (spec 2025-06-18) adr 2025-06-18 2026-10-02 https://raw.githubusercontent.com/modelcontextprotocol/modelcontextprotocol/main/docs/docs/2025-06-18/tutorials/security/security_best_practices.mdx The protocol's own threat page is organised around the confused deputy and forbids token passthrough "To prevent confused deputy attacks, MCP proxy servers MUST implement…" / "'Token passthrough' is an anti-pattern… explicitly forbidden"; "MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server"
15 MCP project Tools page, spec 2025-06-18 adr 2025-06-18 2026-10-02 https://raw.githubusercontent.com/modelcontextprotocol/modelcontextprotocol/main/docs/specification/2025-06-18/server/tools.mdx The spec places a human veto in the loop and declares tool metadata untrusted "there SHOULD always be a human in the loop with the ability to deny tool invocations" / "clients MUST consider tool annotations to be untrusted unless they come from trusted servers"
16 Google "Google's Approach for Secure AI Agents: An Introduction" (Díaz, Kern, Olive) adr 2025-05 2026-10-02 https://storage.googleapis.com/gweb-research2023-media/pubtools/1018686.pdf Google's stated architecture is a hybrid: deterministic runtime policy enforcement as the outer layer, reasoning-based defenses inside "This defense-in-depth approach relies on enforced boundaries around the AI agent's operational environment … acting as guardrails even if the agent's internal reasoning process becomes compromised"; "neither purely rule-based systems nor purely AI-based judgment are sufficient on their own"; "agents must operate under well-defined human control, their powers must be carefully limited according to risk and purpose, and their actions and planning must be observable"
17 Anthropic "Beyond permission prompts: making Claude Code more secure and autonomous" blog 2025-10-20 2026-10-02 https://www.anthropic.com/engineering/claude-code-sandboxing Measured: capability envelopes replaced 84% of per-action approvals; per-action prompts produce approval fatigue "In our internal usage, we've found that sandboxing safely reduces permission prompts by 84%." / "Constantly clicking 'approve' … can lead to 'approval fatigue', where users might not pay close attention to what they're approving"
18 Anthropic sandbox-runtime README source 2025-10 (checked) 2026-10-02 https://github.com/anthropic-experimental/sandbox-runtime The enforcement design: OS primitives, deny-by-default network via host-side proxies, both isolations required "Without file isolation, a compromised process could exfiltrate SSH keys or other sensitive files. Without network isolation, a process could escape the sandbox and gain unrestricted network access." / Linux: "network namespace of the sandboxed process is removed entirely, so all network traffic must go through the proxies running on the host"
19 Anthropic Claude Code sandboxing docs vendor 2026 (checked) 2026-10-02 https://code.claude.com/docs/en/sandboxing The sandbox ships with a model-visible escape hatch (dangerouslyDisableSandbox) whose gate depends on permission mode, and an admin switch to remove it "A sandboxed command has no direct route to the network" / "Claude analyzes the failure and may retry the command with the dangerouslyDisableSandbox parameter" / with allowUnsandboxedCommands: false "the sandbox becomes admin-required. Claude Code then ignores the settings in a repository's files that loosen the sandbox."
20 Google DeepMind / Google Research / ETH Zürich CaMeL code release (paper artifact, arXiv:2503.18813) paper 2025-03 2026-10-02 https://github.com/google-research/camel-prompt-injection The by-design defense extracts control/data flow from the trusted query and runs untrusted data through a quarantined model with no tool access; the authors flag the artifact's own limits "This is a research artifact released to reproduce the results in our paper. The interpreter implementation likely contains bugs … and the implementation might not be fully secure."
21 ETH Zürich SPY Lab AgentDojo (NeurIPS 2024 D&B; paper artifact) paper 2024-06 (paper), docs checked 2026 2026-10-02 https://github.com/ethz-spylab/agentdojo The standard evaluation harness for injection attacks and defenses; adopted by both national AI safety institutes "AgentDojo has been used by US and UK AISI to show the vulnerability of Claude 3.5 Sonnet (new) to prompt injections." (docs/index.md)
22 Meta LlamaFirewall README (PurpleLlama) source 2025-04 (checked 2026) 2026-10-02 https://github.com/meta-llama/PurpleLlama/tree/main/LlamaFirewall Meta's open guardrail layer is scanners composed in a policy engine, including chain-of-thought auditing for goal hijack "AlignmentCheck: A chain-of-thought auditing module that inspects the reasoning process of an LLM agent in real time … to detect goal hijacking, indirect prompt injections, and signs of agent misalignment."
23 OpenAI Agents SDK: human-in-the-loop docs source 2025-2026 (checked) 2026-10-02 https://github.com/openai/openai-agents-python/blob/main/docs/human_in_the_loop.md Approval is a first-class tool property, and the SDK fails closed when it cannot parse what it is approving "Set needs_approval to True to always require approval" / "Callable approval rules fail closed when the SDK cannot safely inspect the arguments."
24 OpenAI Agents SDK: guardrails docs source 2025-2026 (checked) 2026-10-02 https://github.com/openai/openai-agents-python/blob/main/docs/guardrails.md Guardrails are positioned as cost/latency screens with tripwires; parallel execution weakens the guarantee "Blocking execution guarantees that the expensive model does not start; with parallel execution, the expensive model may already have started before the guardrail completes."
25 OpenAI codex-rs core README (sandbox implementation) source 2025-2026 (checked) 2026-10-02 https://github.com/openai/codex/blob/main/codex-rs/core/README.md Codex enforces SandboxPolicy with Seatbelt on macOS and Landlock/bubblewrap on Linux "Network access and filesystem read/write roots are controlled by SandboxPolicy. Seatbelt consumes the resolved policy and enforces it."
26 Google ADK tool-confirmation sample (HITL) source 2025-2026 (checked) 2026-10-02 https://github.com/google/adk-python/blob/main/contributing/samples/hitl/tool_confirmation/README.md Google's agent kit makes per-tool confirmation a framework feature for sensitive actions "It shows how a tool can dynamically request confirmation from the user before proceeding with a sensitive action (e.g., transferring funds)."
27 Hugging Face smolagents: secure code execution source 2025 (checked 2026) 2026-10-02 https://github.com/huggingface/smolagents/blob/main/docs/source/en/tutorials/secure_code_execution.md A framework author's concession that in-process sandboxes cannot be trusted; real isolation means a remote/container executor "no local python sandbox can ever be completely secure … The only way to run LLM-generated code with truly robust security isolation is to use remote execution options like E2B or Docker"
28 GitLab Work item #628875 (Ai::NamespaceSetting cascade) source 2026-09-14 2026-10-02 https://gitlab.com/gitlab-org/gitlab/-/work_items/628875 GitLab ships prompt-injection detection and agent network allow/denylists as namespace settings, and the governance layer has its own bugs: settings do not inherit down the group hierarchy "Prompt injection protection and the Agent Platform network access allowlist/denylist — do not cascade down the group hierarchy. Each group holds its own independent value"
29 ETH Zürich SPY Lab agentdojo on PyPI (release record) source 2024-2026 2026-10-02 https://pypi.org/project/agentdojo/ The benchmark is maintained as an installable package with an active release history PyPI project page, version history (checked 2026-10-02)

Named gaps (primary accounts outside this session's reachable network)

Fetched-link policy: these shaped the public conversation but could not be fetched or re-verified from this session, so the guide names them as context and does not cite them as evidence: Aim Security's EchoLeak write-up (aim.security), Invariant Labs' GitHub MCP exploit post (invariantlabs.ai), Simon Willison's "lethal trifecta" post (simonwillison.net), Meta's "Agents Rule of Two" post (ai.meta.com), the Replit production-database deletion coverage, the AWS security bulletin for the Amazon Q incident (aws.amazon.com), the arXiv PDFs for CaMeL (2503.18813), AgentDojo, and the OpenAI instruction-hierarchy paper, and all conference talks. No conference talk is cited anywhere in this guide for the same reason: zero of the relevant venues (YouTube, USENIX, conference sites) are reachable from this session.