Your model documentation for a customer-facing assistant records the provider's model version it was validated against and the evaluation scores from that run. The Azure deployment of that hosted model is set to auto-update to the provider's default version. Overnight the provider's default changes. What happens next?
Show the full answer Hide the answer
What happens, step by step
- Nothing errors. Microsoft documents that an Azure model deployment set to auto-update moves to the new default within two weeks of the change, and that customers get at least two weeks of notice before a version becomes the default. Requests keep returning 200s, because a model swap is not an error condition anywhere in the stack.
- Your evaluation evidence now describes a model that is no longer serving. The scores in the model card were produced against a version that is no longer behind the endpoint. The document has not changed, so nothing signals this.
- Behaviour shifts in ways your metrics were not built to see. Refusal rate, response length, formatting and tool-call reliability move first. Latency and error rate — the things that are alerted on — often do not move at all.
- The first detection is usually a human. A support agent notices replies got longer, or a downstream parser starts failing on a changed output shape. Median time to notice in this pattern is days, and it is discovered by the complaint rather than the dashboard.
- The audit finding is worse than the behaviour change. An assessor comparing the version in the deployment to the version in the documentation finds they differ, which reads as a control failure across the whole period — even if the new model is better.
Where it amplifies
Anywhere the output is parsed rather than read. A JSON contract held by prompt alone is a coupling to model behaviour that no schema enforces, so a version change propagates into a data pipeline as malformed records rather than as errors.
What stops it
- Pin the version and treat the upgrade as a change. Azure exposes the choice explicitly: upgrade when a new default appears, upgrade only when the current version is retired, or never upgrade automatically. The last option means the deployment stops working at the retirement date. That is the trade-off worth taking for a regulated path: a hard stop you schedule costs an afternoon, and a silent swap costs the validity of every evaluation you have published.
- Use the notice window as the gate. Two weeks is enough to run the evaluation suite against the candidate version in a shadow deployment and compare against the recorded baseline.
- Make the documentation generated, not typed. The model card's validity block — model version, evaluation suite version, run date, dataset snapshot — is emitted by the evaluation job and read from the live deployment. A document that states a version a machine can contradict is the only kind that stays true.
- Alert on the version string, not just on latency. A single check comparing the deployment's reported version against the documented one catches this the morning it happens.
What would have to be true for it to self-heal
The evaluation suite would have to run continuously against production traffic samples with the baseline attached, and the documentation would have to be rebuilt from that run. That is achievable and it is the direction regulation is pushing: the EU AI Act requires technical documentation for high-risk systems under Article 11 and Annex IV to be kept up to date, and instructions for use for deployers under Article 13. The Digital Omnibus regulation that entered into force in July 2026 deferred the Annex III high-risk application date to 2 December 2027, which buys time rather than removing the duty.
When not to pin the version
An internal summarisation tool where a person reads every output, with no parser downstream and no external obligation, is cheaper on auto-update. Pin when the output is consumed by code, when the evaluation result is evidence someone relies on, or when the model's behaviour is part of a contract. Otherwise take the default and re-run the suite on a schedule.