Model Evaluation Service · View 17 of 21 · Operations
Decisions
- A harness version bump is a frame change, so promoting the platform triggers the same re-baseline a rubric edit does.
- A synthetic deliberately-regressed candidate runs on every platform release. A gate that has never blocked anything is not known to work.
- Scorer fixtures are a regression suite for the scorers themselves — a scorer defect returns a plausible wrong number, which is the hardest failure to notice.
Fail closed
- A pipeline that cannot reach the gate does not proceed. Shipping unevaluated to 40 million users is worse than a delayed release, and there is no bypass an outage can trigger.
Assumption
- Platform releases are infrequent relative to candidates, so re-baseline cost from harness bumps is a small share of the budget.