Term Kind Topic What it is
Agent Handoff pattern Multi-Agent Systems The transfer of a task and its context from one specialised agent to another, and the point at which multi-agent systems most often lose information.
Agent Tool Authorisation pattern Tool Calling Enforcing that a tool invoked by a model executes with the requesting user's permissions rather than the application's, and that consequential actions require confirmation.
Agent Transcript Compaction Loop History Summarisation, Working Memory Compaction pattern Agent Architectures Replacing the resolved middle of a long agent loop with a short summary while preserving the original goal and constraints verbatim, so token cost stops growing quadratically and late steps still follow early …
AI Gateway pattern AI-Era Architecture A shared proxy in front of model providers that centralises routing, keys, quotas, caching, logging and safety policy.
AI Gateway Pattern LLM Proxy, Model Router pattern AI Gateways A single control point through which all model calls pass, providing routing, cost control, caching, logging, guardrails and provider abstraction.
Chunked Prefill Prefill Splitting, Interleaved Prefill, Piecewise Prompt Processing pattern AI Cost Management Breaking a long prompt's prefill into pieces interleaved with other sequences' decode steps, so that one enormous request cannot stall every in-flight interactive response.
Contextual Retrieval Contextual Chunk Prefixing, Contextual Embeddings pattern Chunking & Retrieval Prepending a short generated description of where a chunk sits in its document before embedding it, so that a passage full of pronouns and bare figures still matches the query that should find it.
Filtered Vector Search Predicate-Constrained ANN, Metadata Filtering in ANN pattern Vector Databases Combining a metadata predicate with approximate nearest-neighbour search - where the predicate's selectivity, not the corpus size, decides whether results are correct or latency collapses.
Guardrails pattern Guardrails Deterministic checks applied to model inputs and outputs, enforcing constraints that the model itself cannot be relied upon to respect.
Human in the Loop HITL pattern AI-Era Architecture Requiring human review or approval at a defined point in an automated flow, chosen by the reversibility and cost of the action.
Human-in-the-Loop Design pattern Human in the Loop Placing human review at the points where model error is consequential, designed so the review is genuinely effective rather than nominal.
Hybrid Retrieval Dense + Sparse Retrieval, BM25 + Vector pattern RAG Architecture Running lexical keyword search and dense vector search together and fusing the results, because each fails where the other succeeds.
Model Cascade Tiered Inference, Escalation Ladder pattern Model Selection A cheap fast model handling the clear majority of cases with escalation to a larger model or a human for the uncertain ones - usually a large cost reduction with no quality loss, because the expensive path run…
Model Router pattern AI-Era Architecture Directing each request to a model chosen by the task's difficulty, cost and latency budget, rather than sending everything to the largest model available.
Model Routing pattern Model Selection Directing each request to the cheapest model capable of handling it, rather than sending all traffic to the most capable one.
Output Validation Layer pattern Guardrails A deterministic check applied to model output before it is used, treating the model as an untrusted component.
Reranking Cross-Encoder Reranking pattern Reranking Retrieving a wide candidate set cheaply, then reordering it with a more expensive model that scores each candidate against the query directly.
Retrieval Pipeline Stages pattern RAG Architecture The stages that turn a user question into grounded context — query processing, retrieval, reranking and assembly — each independently tunable.
Retrieval-Augmented Generation RAG pattern AI-Era Architecture Retrieving relevant documents at query time and putting them in the model's context, so answers are grounded in your data rather than in training data.
Semantic Cache pattern AI-Era Architecture Caching model responses keyed by the meaning of the request rather than by its exact text, so near-duplicate questions are served without a model call.
Tool Calling Function Calling pattern AI-Era Architecture Giving a model a set of typed function definitions it can request to invoke, with the application executing the call and returning the result.
Tool Result Budget Tool Output Capping, Result Truncation Contract pattern Tool Calling A hard cap on the tokens any single tool may return into the model's context, with pagination and summarisation behind it, so that one unlucky query cannot fill the window and end the run.
Two-Stage Retrieval pattern Reranking Retrieving a broad candidate set cheaply and then reordering it with an expensive, more accurate model.