Turning Principles into Gates That Block
Why AI principles documents change nothing on their own, the properties a gate needs to have force, and the design that keeps a review board from becoming either a bottleneck or a formality.
Nearly every large organisation has published AI principles. Very few can name a project those principles stopped. The gap between the two is the whole of AI governance as an engineering problem, and it is not solved by better principles.
A principle changes behaviour when it is attached to a gate: a point in a process where work cannot proceed until a condition is met, with someone empowered to refuse.
What a gate needs
A defined trigger. Which work requires review, determined by objective criteria rather than by someone's judgement that a project feels significant. Criteria that work in practice are about consequence: does the system make or materially influence decisions about people, does it process sensitive categories of data, is it customer-facing, does it operate autonomously, is it in a regulated domain.
An owner who can say no. A reviewer who can only advise is a consultant. The authority to block, and the organisational backing to survive using it, is what separates a gate from a checkpoint.
Objective criteria. "The model must be fair" cannot be assessed. "Selection rate ratios across the specified groups fall within the stated band, or an exception is documented with reasoning" can be. Subjective criteria produce inconsistent decisions and are the main reason review boards lose credibility.
A bounded timeline. A review with no service level becomes a queue teams route around. Committing to a turnaround, and tiering so low-risk work gets a lightweight path, is what keeps the gate from being bypassed.
An exception path. Some work will proceed despite a finding. Making that explicit, with a named accepter and a recorded rationale, is better than the alternative where exceptions happen informally and leave no record.
Tiering
Uniform process fails in both directions. The workable design classifies systems by consequence and applies proportionate review: self-assessment against a checklist for low-risk internal tools, structured review for systems affecting people, and full review with external input for high-consequence deployments.
Getting the classification right matters more than the depth of the deepest tier, because misclassification is what lets a consequential system through a light path.
When it breaks
Gates arrive too late. A review at the end of a project can only approve or cancel, and cancelling months of work is a decision organisations rarely take, so late gates approve. Engaging at design, when changing course is cheap, is where review changes outcomes.
The board lacks the expertise to evaluate. A committee of senior people without the technical depth to interrogate an evaluation defers to the team presenting. Effective boards include people who can read the evaluation and ask the uncomfortable question.
Nothing is ever refused. A gate with a hundred percent approval rate is either measuring nothing or reviewing work that was never going to fail. The refusal rate is a diagnostic, and a board that has never blocked anything should examine its criteria rather than congratulate itself.
Post-deployment is ungoverned. Approval covers a system as described at review, and systems change: new features, new data, new capabilities added to an agent. Without a trigger for re-review on material change, governance covers a snapshot and the deployed system drifts away from it.
12 flashcards for this concept
Click a card to reveal the answer.