Model Evaluation Service · View 09 of 21 · Data
Decisions
- Redaction happens at admission, before anything is durable. There is no path by which raw production content becomes a dataset example.
- Human labels land in a store separate from scores, so re-scoring the platform's own output can never disturb the ground truth.
Volumes
- 1.05 M scored generations per week → ~55 M score rows and ~2.2 TB of traces per year (stated assumptions, traces at ~40 KB).
- 39,000 judge calls per gate run: 6,500 examples × 3 generations × 2 orderings.
Risks
- The loop back from production is where the erasure obligation enters. See the storage view for why reproducibility is bounded.