Term Kind Topic What it is
Agent Handoff pattern Multi-Agent Systems The transfer of a task and its context from one specialised agent to another, and the point at which multi-agent systems most often lose information.
Agent Loop concept Agent Architectures The cycle in which a model observes state, selects an action, executes a tool and observes the result, repeating until a goal or a limit is reached.
Agent Tool Authorisation pattern Tool Calling Enforcing that a tool invoked by a model executes with the requesting user's permissions rather than the application's, and that consequential actions require confirmation.
Agent Transcript Compaction Loop History Summarisation, Working Memory Compaction pattern Agent Architectures Replacing the resolved middle of a long agent loop with a short summary while preserving the original goal and constraints verbatim, so token cost stops growing quadratically and late steps still follow early …
AI Gateway pattern AI-Era Architecture A shared proxy in front of model providers that centralises routing, keys, quotas, caching, logging and safety policy.
AI Gateway Pattern LLM Proxy, Model Router pattern AI Gateways A single control point through which all model calls pass, providing routing, cost control, caching, logging, guardrails and provider abstraction.
Approximate Nearest Neighbour Index ANN concept Vector Databases An index that trades exactness for speed when finding similar vectors, making large-scale semantic search feasible.
Automation Ratchet Residual Difficulty, Hard-Case Concentration concept Human in the Loop The effect where improving automation makes the cases reaching humans systematically harder, so reviewer throughput falls and error rates rise even as the overall system improves.
Candidate Recall Ceiling First-Stage Recall, Retrieval Ceiling metric Reranking The share of queries whose correct passage appears anywhere in the first-stage candidate set - the hard upper bound on every later stage, since reranking reorders what was retrieved and cannot add to it.
Capability Confinement Least-Privilege Agents, Trust Domain Separation, Action Authorisation concept Prompt Injection Defence Limiting what an agent is authorised to do rather than trying to prevent it from being misled - because a model cannot reliably distinguish instructions from data, so the security boundary must sit outside it.
Chunk Boundary Strategy practice Chunking & Retrieval How source documents are split for embedding, which determines whether retrieved passages are self-contained and coherent.
Chunked Prefill Prefill Splitting, Interleaved Prefill, Piecewise Prompt Processing pattern AI Cost Management Breaking a long prompt's prefill into pieces interleaved with other sequences' decode steps, so that one enormous request cannot stall every in-flight interactive response.
Context Window concept AI-Era Architecture The maximum number of tokens a model can attend to in one request, holding the system prompt, history, retrieved context, tools and the answer.
Context Window Budget concept LLM Application Architecture The finite token allowance per request, treated as an engineering resource to be allocated deliberately between system instructions, retrieved context, history and output.
Contextual Retrieval Contextual Chunk Prefixing, Contextual Embeddings pattern Chunking & Retrieval Prepending a short generated description of where a chunk sits in its document before embedding it, so that a passage full of pronouns and bare figures still matches the query that should find it.
Embedding concept AI-Era Architecture A dense numeric vector representing a piece of content, positioned so that semantically similar content sits nearby.
Embedding Model Migration practice Embeddings The process of moving a corpus to a new embedding model, which requires re-embedding everything because vectors from different models are not comparable.
Embedding Table Collision Hashing Trick Collision, ID Bucket Collision concept Embeddings Two unrelated identifiers mapped to the same row of a fixed-size embedding table, so their learned representations are averaged together and the model quietly treats distinct items as one.
Embeddings concept Embeddings Dense numeric representations of content that place similar things close together — the substrate of semantic search and retrieval.
Escalation Threshold practice Human in the Loop The rule determining when a model's output is acted on automatically and when it is routed to a person.
Evaluation Set Contamination Benchmark Leakage, Golden Set Overfit concept LLM Evaluation The gradual loss of an evaluation set's ability to predict production quality as its cases leak into training data, into prompts, or into months of iteration against the same examples.
Feature Store tool ML Platform A system that computes, stores and serves model input features consistently for both training and inference, eliminating training-serving skew.
Filtered Vector Search Predicate-Constrained ANN, Metadata Filtering in ANN pattern Vector Databases Combining a metadata predicate with approximate nearest-neighbour search - where the predicate's selectivity, not the corpus size, decides whether results are correct or latency collapses.
Guardrail Availability Policy Fail-Open Guardrail Decision, Safety Check Degradation Policy practice Guardrails The decision, made in advance and per action class, about what the system does when a safety check cannot run - because the alternative is that a timeout in a classifier decides your safety posture at 03:00.
Guardrails pattern Guardrails Deterministic checks applied to model inputs and outputs, enforcing constraints that the model itself cannot be relied upon to respect.
Human in the Loop HITL pattern AI-Era Architecture Requiring human review or approval at a defined point in an automated flow, chosen by the reversibility and cost of the action.
Human-in-the-Loop Design pattern Human in the Loop Placing human review at the points where model error is consequential, designed so the review is genuinely effective rather than nominal.
Hybrid Retrieval Dense + Sparse Retrieval, BM25 + Vector pattern RAG Architecture Running lexical keyword search and dense vector search together and fusing the results, because each fails where the other succeeds.
Indirect Prompt Injection concept Prompt Injection Defence An attack in which malicious instructions are placed in content the model will later retrieve, rather than typed by the user.
Inference Cost per Interaction metric AI Cost Management The fully loaded token cost of one user interaction, including retrieval, retries, guardrails and agent loops, measured rather than estimated.
Inference Request Path concept LLM Application Architecture The sequence of stages an LLM application request passes through, each with distinct latency, cost and failure characteristics.
Inference Telemetry practice AI Observability Recording the full context of each model interaction — inputs, outputs, tokens, latency, model version and evaluation scores — so quality and cost can be investigated.
KV Cache Attention Cache, Key-Value Cache, Prefix Cache concept LLM Application Architecture The per-session key and value tensors a transformer must hold in GPU memory to generate each subsequent token - the resource that limits concurrent sessions, and whose reuse across turns is the difference betw…
Least-Privilege Tooling Bounded Tool Permissions practice Prompt Injection Defence Giving each tool the narrowest possible capability and enforcing authorisation at the tool - the decisive control when a model's instructions can be influenced by untrusted content.
LLM Evaluation Evals practice AI-Era Architecture A repeatable measurement of whether an AI system's outputs are good enough, on cases that reflect the actual task.
LLM-as-Judge practice LLM Evaluation Using a language model to score another model's outputs against criteria, making evaluation scalable at the cost of introducing the judge's own biases.
ML Platform concept ML Platform The infrastructure that makes machine learning repeatable — data, features, training, deployment, monitoring — where the model is the small part.
Model Cascade Tiered Inference, Escalation Ladder pattern Model Selection A cheap fast model handling the clear majority of cases with escalation to a larger model or a human for the uncertain ones - usually a large cost reduction with no quality loss, because the expensive path run…
Model Context Protocol MCP protocol AI-Era Architecture An open protocol that standardises how AI applications connect to external tools, data sources and prompts.
Model Drift Monitoring practice AI Observability Detecting that a deployed model's inputs or performance have shifted away from the conditions it was validated under.
Model Router pattern AI-Era Architecture Directing each request to a model chosen by the task's difficulty, cost and latency budget, rather than sending everything to the largest model available.
Model Routing pattern Model Selection Directing each request to the cheapest model capable of handling it, rather than sending all traffic to the most capable one.
Output Validation Layer pattern Guardrails A deterministic check applied to model output before it is used, treating the model as an untrusted component.
Prompt Cache Economics Prefix Cache Break-Even, Cache TTL Selection practice AI Cost Management Deciding what to cache and at which time-to-live by comparing the provider's write premium against its read discount, so a long stable prefix is paid for once rather than on every request.
Prompt Injection concept AI-Era Architecture An attack in which text from an untrusted source is interpreted by the model as instructions rather than as data.
Prompt Injection Defence practice Prompt Injection Defence Defending systems where untrusted content reaches a language model that can take actions — a problem of privilege, not of filtering.
Prompt Registry Prompt Management tool AI-Era Architecture A versioned store of production prompts with their model bindings, parameters and evaluation results, so a prompt change is a reviewable, traceable, reversible deployment.
Prompt Regression Suite practice Prompt & Version Management A set of test cases with expected properties, run against a prompt on every change, to detect quality regressions before deployment.
Prompt Versioning practice Prompt & Version Management Treating prompts as versioned, reviewed, tested and deployable artifacts rather than as strings edited in place.
Quality Proxy Metric Implicit Quality Signal, Behavioural Quality Indicator metric AI Observability A continuously available production signal that moves when answer quality moves - abstention rate, regeneration rate, escalation rate - used to page someone in the hours before a labelled evaluation could ever run.
Reranking Cross-Encoder Reranking pattern Reranking Retrieving a wide candidate set cheaply, then reordering it with a more expensive model that scores each candidate against the query directly.
Retrieval Evaluation practice LLM Evaluation Measuring whether the right context was retrieved, separately from whether the answer was good, because the two failures need different fixes.
Retrieval Grounding Grounded Generation concept RAG Architecture Constraining a model's answer to content retrieved from an authoritative corpus, with citations and an abstention path - so that output quality becomes a retrieval problem rather than a model problem.
Retrieval Pipeline Stages pattern RAG Architecture The stages that turn a user question into grounded context — query processing, retrieval, reranking and assembly — each independently tunable.
Retrieval-Augmented Generation RAG pattern AI-Era Architecture Retrieving relevant documents at query time and putting them in the model's context, so answers are grounded in your data rather than in training data.
Retrieval-Generation Separation Evaluate Retrieval Independently, Two-Stage Debugging practice RAG Architecture Measuring whether the correct passage was retrieved, separately from whether the answer was correct - the single diagnostic that turns unfalsifiable RAG debugging into a specific measurable defect.
Review Queue Backpressure Human Queue Saturation, Reviewer Capacity Limit concept Human in the Loop What a human-review step does when cases arrive faster than reviewers can clear them - a decision that must be made per action class in advance, because the alternative is an unplanned auto-approval or an unbo…
Semantic Cache pattern AI-Era Architecture Caching model responses keyed by the meaning of the request rather than by its exact text, so near-duplicate questions are served without a model call.
Semantic Chunking practice Chunking & Retrieval Splitting documents along their meaning and structure rather than at fixed character counts, because retrieval quality is bounded by chunk quality.
Serving Path Divergence Per-Pool Quality Drift, Heterogeneous Fleet Skew concept AI Observability One configuration in a fleet of otherwise identical inference paths behaving differently from its peers - detectable by comparing paths against each other, and invisible to any metric averaged across them.