concept

Validated Envelope

also called Model Use Envelope, Validation Boundary

The explicit conditions a model's validation actually covers - input schema, population, feature sources, thresholds, upstream versions - outside which the sign-off does not hold and the serving path must refuse or escalate.

model-riskvalidationpromotion-gatetraining-serving-skewdrift

Validation of a credit model takes six weeks and produces a signed report. The model retrains every week. By the time the report is filed the weights it describes have been replaced, and they will be replaced 51 more times before the next review. The sign-off is attached to the wrong object.

The second, more common version has nothing to do with retraining. A model validated on one country's loan book is applied to a newly acquired portfolio with a different applicant mix. Nothing in the model changed, monitoring is green, and the validation no longer covers what is running.

Why it matters

Model risk is not only the risk that a model is poorly built. It is equally the risk that a sound model is used outside the conditions it was built for, which is why supervisory model risk guidance in US banking (2011) treats use outside intended purpose as a first-class source of risk and requires ongoing monitoring rather than one-time approval.

An envelope makes "outside the conditions" machine-checkable. Without one, the boundary of the sign-off exists only in the validator's head, so the first person to point the model at a new segment is making a risk decision without knowing it.

Implementation patterns

  • Write the envelope as a contract: input schema with types and ranges, feature source ids and owning teams, the population the training data represents, the acceptance thresholds a candidate must clear, and pinned versions of anything upstream - hosted model, embedding model, transformation library.
  • A promotion gate inside the retraining pipeline that evaluates each candidate on a frozen held-out slice and refuses promotion if a threshold is missed or the envelope has changed.
  • Runtime assertions in serving: per-feature range and null-rate checks plus a population membership test, with a defined out-of-envelope behaviour - fall back to a rules path, escalate, or refuse. Range checks cost well under 1 ms; a membership model costs another inference, and 20 ms of added latency is cheap next to a decision you cannot defend.
  • Name the computation, not just the field. A feature computed one way in training and another in serving passes any schema check and is the most common silent model defect.
  • Stamp the envelope version on every prediction, so a breach lets you select exactly which decisions to revisit instead of re-reviewing a year.
  • Re-validate when the envelope changes, not on a calendar.

Industry example

The shape is clearest in lending. An underwriting model is validated against one book, its envelope implicitly set by that book's income distribution and product mix. The business then launches a channel whose applicants are younger and thinner-file. Approval rates move, the offline metric looks fine because it is computed on logged data from that same channel, and the first real signal is default performance 9 to 18 months later. A population membership check raises it on day one, as a configuration question rather than a credit loss.

Failure scenarios

  • An upstream team changes a feature's units or backfills history. Every weekly model clears its own thresholds; the input distribution check is the only thing that notices.
  • A hosted embedding model version changes, so stored and query vectors no longer share a space and retrieval quality collapses with no error anywhere.
  • An envelope that exists only in the validation document, so it is documentation rather than a control.

Trade-offs

Runtime assertions add latency and generate false alarms, and a tight envelope blocks legitimate expansion, producing waivers against the model's own governance. A loose envelope passes everything and proves nothing. The defensible position is tight on what is cheap to check and genuinely breaks the model (schema, ranges, null rates, upstream versions) and explicit rather than automated on population, where the judgement is a business one.

When not to use it

For a model retrained yearly by hand, the artefact is the right unit and the envelope machinery costs more than it returns. Where no consequential decision is attached, registration and monitoring are proportionate, and where the feature set is three fields from one owner the schema is the contract.

Interview question

Q: An independent validator signs off a model that retrains weekly. Six months later a regulator asks what exactly was approved. What do you want to show, and what has to be in the pipeline for that answer to exist?

What a strong answer covers: that the validated object is the pipeline plus an explicit envelope rather than the weights; the envelope's fields; a promotion gate enforcing thresholds; runtime assertions and their out-of-envelope behaviour; the envelope version stamped on each prediction so affected decisions are selectable; and re-validation triggered by envelope change rather than by date.

Quick check

Quiz: A model is validated in six weeks and retrains weekly. What did validation approve? — The training pipeline and the envelope of models it may produce: schema, population, feature sources, thresholds and pinned upstream versions. Not the weights, which are gone by the time the report is signed.

Flashcard: Which single check catches an upstream team changing a feature's units? — A runtime assertion on that feature's range and null rate against the validated envelope, because the model's own offline thresholds are computed on the changed data and look fine.