Model Evaluation Service  ·  View 01 of 21  ·  Context and scope

System Context

What the platform is asked, by whom, and which systems it must reach to answer.

Editable source SVG draw.io All views
Release path Deployment Pipeline Assistant Serving People ML Engineer Quality Lead Human Rater Release On-Call Inference surfaces Vertex AI External Providers Model Evaluation Service Evidence · verdict · gate Signals and governance Production Telemetry Privacy Controls submits candidates owns the gates labels · adjudicates acts on verdicts candidate endpoints judge · rival models asks for a verdict the incumbent traces · guardrails consent · erasure Model Evaluation Service — System Context External / third party Person or role Security / platform synchronous batch The platform answers one question for the pipeline: is this candidate better than the incumbent, and where is it worse. v 1.0 · owner Data & AI Global Practice · date 2026-09

Decisions

  • The platform answers one question: is this candidate better than the incumbent, and on which slices is it worse.
  • The deployment pipeline is a consumer, not an owner — it asks for a verdict and is bound by a fail-closed contract.
  • Production telemetry is an input. The platform is not an observability product and does not serve live dashboards.

Assumptions

  • 40 million monthly active users, 25 release candidates per week (stated assumption).
  • One primary in-house model on Vertex AI plus two external provider models, one of which also serves as the judge.

Out of scope

  • Training, fine-tuning and serving. The platform measures a candidate; it does not produce one.
  • Choosing which candidate to build — that is a research decision this platform informs.