Model Evaluation Service  ·  View 09 of 21  ·  Data

Data Flow

From a user's conversation to a gate-bearing example, and from a candidate's output to a verdict.

Editable source SVG draw.io All views
Sources Production Traces 40 M MAU Curated Examples Incident Reports Admission Redaction Sensitive Data Protection Provenance Stamp consent · basis Corpus Frozen Golden Set immutable Rolling Set refreshed Targeted Sets from incidents Execution Harness 3 gens / example Run Traces ~2.2 TB / year Scoring Scorers + Judges Human Labels separate store Facts Score Store 55 M rows / year Verdict Record 7 years Consumers Slice Reports Offline-Online Correlation Model Evaluation Service — Data Flow External / third party Security / platform Application we own Data store Every score carries its frame. The loop back from production into targeted sets is the only path by which user data becomes an example, and it passes through redaction before anything is durable. v 1.0 · owner Data & AI Global Practice · date 2026-09

Decisions

  • Redaction happens at admission, before anything is durable. There is no path by which raw production content becomes a dataset example.
  • Human labels land in a store separate from scores, so re-scoring the platform's own output can never disturb the ground truth.

Volumes

  • 1.05 M scored generations per week → ~55 M score rows and ~2.2 TB of traces per year (stated assumptions, traces at ~40 KB).
  • 39,000 judge calls per gate run: 6,500 examples × 3 generations × 2 orderings.

Risks

  • The loop back from production is where the erasure obligation enters. See the storage view for why reproducibility is bounded.