Model Evaluation Service  ·  View 03 of 21  ·  People and journeys

Actors and Their Core Journeys

Six kinds of people, what each of them is actually trying to achieve, and the end user who must never meet the platform.

Editable source SVG draw.io All views
The people who ship models ML Engineer 25 candidates / week Goal — Find out whether my change made the assistant better, before I have to defend it in a review. Core journeys Run a smoke check ≤ 8 min, waits for it Take a candidate through the gate ≤ 90 min Argue with a block wants the examples Product Manager owns the assistant Goal — Know what this release actually changes for users, in plain terms, before it goes out. Core journeys Read the release evidence Check a cohort that matters language · task The people who own quality Quality Lead 18 suites · 34 gated slices Goal — Keep the gate trusted — sensitive enough to catch regressions, honest enough that nobody routes around it. Core journeys Author a suite and its gate Approve a frame change costs a re-baseline Review the waiver register Human Rater 40 reviewer-h / week Goal — Judge which answer is actually better, without being told which one the company hopes wins. Core journeys Label the calibration set Adjudicate a blocked regression blinded pairwise The people who carry the release Release On-Call rollout + rollback Goal — Ship when the evidence says ship, and get the old model back fast when it does not. Core journeys Watch a canary Roll back on a guardrail breach ≤ 5 min Privacy Officer governs production-derived data Goal — Be able to say exactly which user data is in which dataset, and get it out when asked. Core journeys Trace an example to its consent Action a deletion request Who this is not for End User 40 M monthly Goal — Never meet this platform, and never meet the release it stopped. Core journeys Gets a better assistant or does not get a worse one Model Evaluation Service — Actors and Their Core Journeys The end user is the beneficiary and never a participant — the only evidence they generate is telemetry, redacted before it becomes an example. v 1.0 · owner Data & AI Global Practice · date 2026-09

What this lands

  • The ML engineer and the quality lead want opposite things from the same gate: speed and strictness. Most of the design tension lives here.
  • The human rater is an instrument the platform is calibrated against, which is why blinding is a requirement and not a courtesy.
  • The privacy officer is an actor because production-derived examples give the platform a lawful-basis obligation.

Assumptions

  • 40 reviewer-hours per week of human capacity — the one input that does not scale with compute (stated assumption).
  • 18 suites and 34 declared gate-bearing slices.

Deliberate omission

  • The end user has one journey and no touchpoint. They benefit from a release that is stopped, and never know it happened.