AI Executive Office — CXO Assistant Platform  ·  View 24 of 30  ·  6 · Operations

Model and Prompt Lifecycle

The loop that keeps answers honest as models, prompts and the business all change underneath them.

Editable source SVG draw.io All views
Curate golden questions per tenant Evaluate offline accuracy · grounding · refusal Gate no regression, no release Canary one tenant, one ring Observe live grounding and feedback Mine failures abstentions and thumbs-down Prompt and model registry versioned set scores per class promote or stop real questions failure classes new golden cases Model and Prompt Lifecycle — the Loop That Keeps Answers Honest Data store Application we own Decision point Security / platform Every model or prompt change re-enters at Curate. A tenant may pin a version and opt in to upgrades on its own schedule. v 1.0 · owner Data & AI Global Practice · date 2026-09

The decision

  • Evaluation is a release gate built before launch, not an assurance activity added afterwards. Retrieval and generation are scored separately so a regression has a stage rather than a shrug
  • Golden questions are curated per tenant, because a ministry's questions and a bank's questions fail differently. The customer writes the go-live set
  • Production failures — abstentions, rejections, corrections — are mined back into the golden set. The loop's real input is the platform's own mistakes

Assumptions

  • Model deprecation is assumed, not hoped against. A tenant may pin a version and opt in to upgrades on its own schedule, within a supported window
  • Scoring uses a model of a different family from the one being scored where the judgement is automated, plus human review on a sample

Risks

  • Evaluation cost is real and recurring. It is budgeted per release rather than absorbed, because the first thing a cost-reduction exercise cuts is the gate that protects quality