practice

LLM-as-Judge

Using a language model to score another model's outputs against criteria, making evaluation scalable at the cost of introducing the judge's own biases.

evaluationllmtesting

Human evaluation is the gold standard and does not scale to the cadence at which prompts, models and retrieval configurations change. Automated metrics based on string overlap are cheap and largely uninformative for open-ended generation. LLM-as-judge sits between: a model scores outputs against a rubric, at volume, for a few cents each.

It works well enough to be the industry default for regression testing, and it has documented biases that must be controlled rather than ignored. Judges prefer longer answers. They prefer their own family's outputs. They are sensitive to presentation order in pairwise comparison. And they are poorly calibrated on absolute scales, drifting between runs.

The practices that make it trustworthy. Prefer pairwise comparison over absolute scoring, since "which of these two is better" is a far more stable judgement than "rate this 1 to 10". Randomise order and run both directions to cancel position bias. Use a specific rubric with concrete criteria rather than asking for overall quality. Ask for reasoning before the verdict. And, critically, validate the judge against human labels on a sample — a judge that agrees with human raters 60% of the time is measuring something other than what you intended, and you cannot know without checking.

The deployment shape that works: a curated evaluation set of real queries with expected properties, run automatically on every prompt or model change, with results tracked over time so a regression is visible before release rather than discovered by users.