Model Evaluation Service · View 01 of 21 · Context and scope
Decisions
- The platform answers one question: is this candidate better than the incumbent, and on which slices is it worse.
- The deployment pipeline is a consumer, not an owner — it asks for a verdict and is bound by a fail-closed contract.
- Production telemetry is an input. The platform is not an observability product and does not serve live dashboards.
Assumptions
- 40 million monthly active users, 25 release candidates per week (stated assumption).
- One primary in-house model on Vertex AI plus two external provider models, one of which also serves as the judge.
Out of scope
- Training, fine-tuning and serving. The platform measures a candidate; it does not produce one.
- Choosing which candidate to build — that is a research decision this platform informs.