Evidence a Regulator Will Accept
The difference between a policy and evidence of its operation, what an assessor actually asks for, and how to instrument a system so compliance artefacts are produced automatically rather than assembled retrospectively.
Compliance programmes reliably produce policies and reliably fail to produce evidence. The distinction matters because an assessor's question is never "do you have a policy on this"; it is "show me that this happened, for this system, on this date, decided by this person".
What evidence looks like
A policy states an intention. It is necessary and it establishes nothing about what occurred.
A record shows a specific instance: this assessment, completed on this date, by this person, with this conclusion. A policy requiring impact assessments plus five completed assessments for five deployed systems is evidence; the policy alone is a statement of intent.
A control with an audit trail is stronger still. A deployment pipeline that refuses unapproved models produces evidence continuously, as a byproduct of operating, and cannot be retroactively fabricated. This is the form of evidence assessors weight most heavily, because it is the hardest to construct after the fact.
The pattern that follows is to instrument the process so artefacts are emitted automatically. Approvals recorded in a system rather than in email. Evaluation results written by the pipeline rather than pasted into a document. Model versions and their approval status linked in the registry. Data lineage generated rather than drawn.
What the AI Act asks for specifically
For a high-risk system: technical documentation covering design, development, and the risk management system; a record of the data governance measures applied to training, validation and testing data; logs the system generates automatically and their retention; the human oversight measures and how they are implemented; accuracy, robustness and cybersecurity measures with the metrics used; the conformity assessment; and post-market monitoring.
For GPAI models: technical documentation, information for downstream providers, a copyright policy, and the public summary of training content.
The recurring theme is that each item names an artefact rather than a property, so the compliance question is whether the artefact exists and is current, not whether the underlying property holds.
When it breaks
Retrospective assembly is detectable and expensive. A programme that generates documentation in the weeks before an assessment produces documents that are internally consistent and disconnected from the systems they describe, which an assessor with technical understanding notices. It is also several times more expensive than continuous generation.
Documentation drifts from the system. A technical file describing version 3 of a model that has been retrained eleven times since is inaccurate, and inaccuracy in a compliance artefact is worse than a gap because it is a false statement rather than a missing one. Regeneration on change is what keeps them aligned.
Nobody owns the artefact. Documentation without a named owner and a review cadence goes stale by default. This is the same failure as an orphaned dataset in a catalogue, with a regulator attached.
Evidence and effectiveness are different claims. A complete technical file for a system that performs badly for a subgroup is complete documentation of a problem. Compliance evidence establishes that the process ran, and the substantive question of whether the system is good enough is a separate one that the paperwork can make visible but not answer.
10 flashcards for this concept
Click a card to reveal the answer.