Term Kind Topic What it is
Chunk Boundary Strategy practice Chunking & Retrieval How source documents are split for embedding, which determines whether retrieved passages are self-contained and coherent.
Embedding Model Migration practice Embeddings The process of moving a corpus to a new embedding model, which requires re-embedding everything because vectors from different models are not comparable.
Escalation Threshold practice Human in the Loop The rule determining when a model's output is acted on automatically and when it is routed to a person.
Guardrail Availability Policy Fail-Open Guardrail Decision, Safety Check Degradation Policy practice Guardrails The decision, made in advance and per action class, about what the system does when a safety check cannot run - because the alternative is that a timeout in a classifier decides your safety posture at 03:00.
Inference Telemetry practice AI Observability Recording the full context of each model interaction — inputs, outputs, tokens, latency, model version and evaluation scores — so quality and cost can be investigated.
Least-Privilege Tooling Bounded Tool Permissions practice Prompt Injection Defence Giving each tool the narrowest possible capability and enforcing authorisation at the tool - the decisive control when a model's instructions can be influenced by untrusted content.
LLM Evaluation Evals practice AI-Era Architecture A repeatable measurement of whether an AI system's outputs are good enough, on cases that reflect the actual task.
LLM-as-Judge practice LLM Evaluation Using a language model to score another model's outputs against criteria, making evaluation scalable at the cost of introducing the judge's own biases.
Model Drift Monitoring practice AI Observability Detecting that a deployed model's inputs or performance have shifted away from the conditions it was validated under.
Prompt Injection Defence practice Prompt Injection Defence Defending systems where untrusted content reaches a language model that can take actions — a problem of privilege, not of filtering.
Prompt Regression Suite practice Prompt & Version Management A set of test cases with expected properties, run against a prompt on every change, to detect quality regressions before deployment.
Prompt Versioning practice Prompt & Version Management Treating prompts as versioned, reviewed, tested and deployable artifacts rather than as strings edited in place.
Retrieval Evaluation practice LLM Evaluation Measuring whether the right context was retrieved, separately from whether the answer was good, because the two failures need different fixes.
Retrieval-Generation Separation Evaluate Retrieval Independently, Two-Stage Debugging practice RAG Architecture Measuring whether the correct passage was retrieved, separately from whether the answer was correct - the single diagnostic that turns unfalsifiable RAG debugging into a specific measurable defect.
Semantic Chunking practice Chunking & Retrieval Splitting documents along their meaning and structure rather than at fixed character counts, because retrieval quality is bounded by chunk quality.
Small Model Routing Model Cascade, Tiered Inference practice Model Selection Sending each request to the smallest model that can handle it, escalating to a larger one only when needed.
Token Budget Enforcement practice AI Gateways Limiting token consumption per user, tenant, feature or time window at a central point, so cost cannot run away unobserved.
Token Cost Attribution practice AI Cost Management Assigning inference spend to features, tenants and users, so that cost can be managed by the people who influence it.
Tool Schema Design practice Tool Calling Defining the tools available to a model — names, descriptions, parameters and errors — in a way that makes correct selection likely.