AI Assurance & Audit advanced 7 min read 10 flashcards

Assurance for Continuously Changing Systems

Why point-in-time assurance is a poor fit for systems that retrain weekly, what continuous assurance requires instead, and how to define the change that resets the conclusion.

Audit and certification were built for artefacts that change slowly: a financial statement covering a closed period, a product that ships and stays shipped. A model retrained weekly, with prompts edited daily and a provider updating the underlying model without notice, has no closed period. Assurance has to change shape, and the changes are specific.

From conclusions to controls

The move that makes assurance tractable is to shift the object from the artefact to the process that produces it.

Rather than asserting that this model is fair, the assurance asserts that every model deployed passes a fairness evaluation against a defined threshold, that the evaluation is automated and cannot be bypassed, that failures block deployment, and that exceptions are recorded and approved. That claim survives retraining, because it is a claim about the pipeline.

This is the same logic that makes continuous controls monitoring work in financial systems, and it has the same requirement: the control must be automated and produce evidence, because a manual control cannot be asserted continuously.

What has to exist

Automated gates that block. A control that warns is not assurable, because the assurance would have to cover every decision to proceed anyway.

Evidence generated per deployment, not per audit. Every release emits its evaluation results, approvals and versions, so the evidence for any point in the period exists without anyone having assembled it.

A defined change taxonomy. Not every change resets the conclusion. A scheduled retrain on the same data pipeline with results within the established band is routine; a new data source, a new capability, a change of base model, or a threshold change is material and triggers re-assessment. Writing that taxonomy down is what makes continuous assurance bounded rather than infinite.

Monitoring as evidence. Production metrics demonstrating that behaviour stayed within expectations across the period substantiate the claim between deployments, which point-in-time testing cannot.

When it breaks

Provider-side changes are outside the pipeline. A hosted model updated behind a stable name changes the system with no deployment and no gate. The only controls available are pinning to dated versions where offered, and running a fixed evaluation set continuously so the change is detected, and both need to be in the assurance scope explicitly.

The change taxonomy is written optimistically. Classifying changes as routine is the path of least resistance, and a taxonomy that makes everything routine has removed the trigger. Independent review of the classification, as with risk tiering, is the correction.

Continuous evidence is voluminous and unread. Generating an artefact per deployment produces thousands per year, which nobody reviews. The assurance value comes from sampling them and from the gates having blocked things, so a record of what was blocked is more informative than the record of what passed.

Assurance frameworks have not caught up. Most certification regimes still issue point-in-time certificates with periodic surveillance, which fits a weekly-changing system poorly. Supplementing a certificate with continuous control evidence is currently the practical arrangement, and it is an addition to the framework rather than something it asks for.

Check yourself

10 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track