LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
4 to work through
-
advanced
A platform ships AI features and cannot tell whether changes improve or degrade quality. What evaluation infrastructure is required, and in what order?
2 min answer -
advanced
A team ships an LLM feature and cannot tell whether changes improve it. What evaluation infrastructure is needed, and what does it not solve?
2 min answer -
advanced
You are asked to prove an AI assistant is good enough to launch. How do you construct the evidence?
2 min answer -
advanced
Your RAG assistant gives confident answers that are subtly wrong. Where do you look first?
2 min answer
2 terms in this topic
LLM-as-Judge
Using a language model to score another model's outputs against criteria, making evaluation scalable at the cost of introducing the judge's own biases.
practiceRetrieval Evaluation
Measuring whether the right context was retrieved, separately from whether the answer was good, because the two failures need different fixes.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.