intermediate 3 min answer

Your security design review template asks for data flows, trust boundaries, authentication and encryption. A team submits a support assistant that reads customer ticket text and calls three internal APIs with a service account, one of which issues refunds. The template produces no findings. What would you add to it, and what would you leave alone?

security-design-reviewprompt-injectionconfused-deputyagentstrust-boundaries
Show the full answer Hide the answer

What is actually required

The template is not wrong, it is incomplete in one specific way: it assumes the thing that decides what call to make is your code. Every question it asks (who authenticates, where the boundary is, what is encrypted) still applies and still earns its place. What it has no question for is the component that turns attacker- controlled text into a privileged action.

A customer writes the ticket body. The ticket body reaches a model. The model chooses which of three tools to invoke. One of those tools moves money. Untrusted input now sits upstream of a credential that the untrusted party does not hold, which is the confused deputy, restated. The template's data-flow question records that tickets flow into the assistant and that the assistant calls the refund API. It has no way to record that one causes the other.

The one change that matters

Add a tool surface section, and make it enumerate, per tool: the action, whose authority it executes under, the per-call and per-period bound, and whether a human confirms the specific effect before it lands.

The engineering consequences follow directly:

  • The service account is the defect. Each tool call should carry a token derived from the requesting human's session, so the assistant can never do more than the agent operating it could do by hand. Where there is no human in the request (a batch triage job), the token is scoped to one tool and one tenant.
  • Limits live in the tool, not in the prompt. The refund endpoint enforces its own cap, for example one refund per ticket and a ceiling per account over 24 hours, because a prompt instruction is a suggestion to a probabilistic system and a server-side check is not. Without the server-side ceiling the blast radius of one crafted ticket is every refund the service account can issue, which is the failure mode that costs money.
  • Idempotency keys are generated by the caller, not the model, or a retried turn refunds twice.
  • The output path is a boundary too. If the assistant's text reaches the customer unread, retrieved internal content can be exfiltrated by instruction.

What I would leave alone

Do not bolt on a forty-question AI annex. The trust boundary and data-flow questions carry most of the value and teams already answer them; adding a parallel questionnaire splits attention and gets the shallower answer from both. One section, four fields per tool.

When not to add it

Prefer the existing template unchanged for an assistant that only retrieves and summarises, with no side-effecting tool and no path to a customer without a human pressing send; there the new section is a formality and the extra fields cost review time you do not have. The trigger that makes deep review mandatory is a model output reaching a side-effecting call or an external recipient without a person confirming that specific effect. Write that trigger down, because a review function of six people supporting two hundred engineers cannot read everything, and in production the difference between a reviewed and an unreviewed agent is whether that one sentence was in the trigger list.