AI Executive Office — CXO Assistant Platform · View 24 of 30 · 6 · Operations
The decision
- Evaluation is a release gate built before launch, not an assurance activity added afterwards. Retrieval and generation are scored separately so a regression has a stage rather than a shrug
- Golden questions are curated per tenant, because a ministry's questions and a bank's questions fail differently. The customer writes the go-live set
- Production failures — abstentions, rejections, corrections — are mined back into the golden set. The loop's real input is the platform's own mistakes
Assumptions
- Model deprecation is assumed, not hoped against. A tenant may pin a version and opt in to upgrades on its own schedule, within a supported window
- Scoring uses a model of a different family from the one being scored where the judgement is automated, plus human review on a sample
Risks
- Evaluation cost is real and recurring. It is budgeted per release rather than absorbed, because the first thing a cost-reduction exercise cuts is the gate that protects quality