Term Kind Topic What it is
Agent Loop concept Agent Architectures The cycle in which a model observes state, selects an action, executes a tool and observes the result, repeating until a goal or a limit is reached.
Approximate Nearest Neighbour Index ANN concept Vector Databases An index that trades exactness for speed when finding similar vectors, making large-scale semantic search feasible.
Automation Ratchet Residual Difficulty, Hard-Case Concentration concept Human in the Loop The effect where improving automation makes the cases reaching humans systematically harder, so reviewer throughput falls and error rates rise even as the overall system improves.
Capability Confinement Least-Privilege Agents, Trust Domain Separation, Action Authorisation concept Prompt Injection Defence Limiting what an agent is authorised to do rather than trying to prevent it from being misled - because a model cannot reliably distinguish instructions from data, so the security boundary must sit outside it.
Context Window concept AI-Era Architecture The maximum number of tokens a model can attend to in one request, holding the system prompt, history, retrieved context, tools and the answer.
Context Window Budget concept LLM Application Architecture The finite token allowance per request, treated as an engineering resource to be allocated deliberately between system instructions, retrieved context, history and output.
Embedding concept AI-Era Architecture A dense numeric vector representing a piece of content, positioned so that semantically similar content sits nearby.
Embedding Table Collision Hashing Trick Collision, ID Bucket Collision concept Embeddings Two unrelated identifiers mapped to the same row of a fixed-size embedding table, so their learned representations are averaged together and the model quietly treats distinct items as one.
Embeddings concept Embeddings Dense numeric representations of content that place similar things close together — the substrate of semantic search and retrieval.
Indirect Prompt Injection concept Prompt Injection Defence An attack in which malicious instructions are placed in content the model will later retrieve, rather than typed by the user.
Inference Request Path concept LLM Application Architecture The sequence of stages an LLM application request passes through, each with distinct latency, cost and failure characteristics.
KV Cache Attention Cache, Key-Value Cache, Prefix Cache concept LLM Application Architecture The per-session key and value tensors a transformer must hold in GPU memory to generate each subsequent token - the resource that limits concurrent sessions, and whose reuse across turns is the difference betw…
ML Platform concept ML Platform The infrastructure that makes machine learning repeatable — data, features, training, deployment, monitoring — where the model is the small part.
Prompt Injection concept AI-Era Architecture An attack in which text from an untrusted source is interpreted by the model as instructions rather than as data.
Retrieval Grounding Grounded Generation concept RAG Architecture Constraining a model's answer to content retrieved from an authoritative corpus, with citations and an abstention path - so that output quality becomes a retrieval problem rather than a model problem.
Serving Path Divergence Per-Pool Quality Drift, Heterogeneous Fleet Skew concept AI Observability One configuration in a fleet of otherwise identical inference paths behaving differently from its peers - detectable by comparing paths against each other, and invisible to any metric averaged across them.
Subagent Context Isolation Parallel Context Windows, Context Partitioning Across Agents concept Multi-Agent Systems The property that makes multi-agent systems worth their cost - each subagent explores with its own context window and returns only a condensed result, so the system reads far more than one context could hold.
Tool Authorisation Boundary Model as Untrusted Proposer, Authorise Outside the Model concept Prompt Injection Defence Authorising every tool invocation against the initiating user's own permissions, outside the model - because the model cannot distinguish instructions from data and no prompt-level defence is reliable.
Training-Serving Skew Online-Offline Skew, Feature Skew concept ML Platform A divergence between the features a model was trained on and the features computed at serving time - producing a model that performs well offline and worse in production, with nothing erroring.