Search the practice set
126 questions, 454 terms and 400 topics in 20 areas.
8 results for “Research & Evaluation”
LLM Evaluation
A repeatable measurement of whether an AI system's outputs are good enough, on cases that reflect the actual task.
Architecture Trade-off Analysis Method
A structured evaluation that scores an architecture against prioritised quality-attribute scenarios and identifies the points where those attributes conflict.
Prompt Registry
A versioned store of production prompts with their model bindings, parameters and evaluation results, so a prompt change is a reviewable, traceable, reversible deployment.
Prompt Versioning
Treating prompts as versioned, reviewed, tested artefacts rather than as strings edited in place.
Research & Evaluation
Assessing a technology quickly without adopting it by accident.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
Model Selection
Capability, latency, cost and the evaluation that decides between them.