The controls are structural, not lexical. A filter searching for the phrase ignore previous instructions has already lost — the evidence slot is never read as instruction in the first place.
The renderer, not the model, decides what becomes a link. Only cited sources become clickable, which closes exfiltration through a model-authored URL completely.
The tool broker checks provenance: a tool request whose origin is a document body is rejected before any policy question is even asked.
Residual risk, stated
A legitimate external page can still be wrong and still be cited. Authority weighting and visible citation make that a reviewable outcome rather than an invisible one.
Absence of results leaks a little information about what exists. Accepted and documented, because the alternative — uniform fake results — is worse for every honest user.
Over-refusal is the cost of these controls, and false refusal rate is gated in view 32 so it cannot creep.
Assurance
220 adversarial cases run on every build, including documents with white text, hostile alt text, hostile slide notes and hostile spreadsheet comments.
A red-team exercise twice a year, with findings feeding the tool and source policies in view 40.