Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
4 to work through
-
advanced
An AI assistant reads code and issues from repositories, including untrusted ones, and can call tools. Why is prompt injection a structural problem, and what actually mitigates it?
2 min answer -
advanced
An agent reads customer emails and can issue refunds. What is the threat and how do you bound it?
2 min answer -
advanced
An agent with tool access processes untrusted content. Why is prompt injection not solvable at the prompt layer, and what architectural controls actually limit the damage?
3 min answer -
advanced
An application passes user input into a model that can call tools. What is the threat, and where must the control live?
2 min answer
5 terms in this topic
Capability Confinement
Limiting what an agent is authorised to do rather than trying to prevent it from being misled - because a model cannot reliably distinguish instructi…
conceptIndirect Prompt Injection
An attack in which malicious instructions are placed in content the model will later retrieve, rather than typed by the user.
practiceLeast-Privilege Tooling
Giving each tool the narrowest possible capability and enforcing authorisation at the tool - the decisive control when a model's instructions can be …
practicePrompt Injection Defence
Defending systems where untrusted content reaches a language model that can take actions — a problem of privilege, not of filtering.
conceptTool Authorisation Boundary
Authorising every tool invocation against the initiating user's own permissions, outside the model - because the model cannot distinguish instruction…
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.