Model Evaluation Service  ·  View 14 of 21  ·  Runtime

Scoring by Family

Five lanes, what each costs, and which of them is allowed to stop a release.

Editable source SVG draw.io All views
Trigger Input Method Output Authority Deterministic Every example Output + expectation Schema · tools · regex Pass / fail Blocking Model judge Every example Pair + rubric Pinned judge · 2 orders Score + rationale Blocking if κ ≥ 0.75 Calibration monitor Continuous Fixed calibration set Re-score · compare κ Drift alarm Retires a frame Human audit Sampled fraction Blinded pair Rater preference Label Ground truth Human adjudication Every BLOCK Moved examples Blinded review Upheld / overturned Final on a refusal Scoring by Family — What Each Kind of Scorer Costs and Proves Application we own Data store Decision point External / third party Security / platform Risk / gap Interface / broker Person or role The judge is the only lane whose authority is conditional. Below 0.65 κ against the calibration set it may still produce scores, but those scores are observational and cannot block a release. v 1.0 · owner Data & AI Global Practice · date 2026-09

Decisions

  • The judge is the only lane whose authority is conditional: below 0.65 κ it still scores, but observationally — it cannot block.
  • Human adjudication is final on a refusal. A machine may raise a block; a person confirms one.
  • The calibration monitor can retire a frame. Silent judge drift is caught here or it is not caught at all.

Assumptions

  • Judge-human agreement ≥ 0.75 Cohen's κ to gate, ≥ 0.65 to be admissible at all; calibration set ≥ 1,200 labelled examples per rubric (stated assumptions).
  • 40 reviewer-hours per week across audit and adjudication.

Open

  • Where human judgment should sit — calibration, adjudication, or continuous audit — is Core Architecture Question 3. This view shows all three, which is the expensive answer.