A claims system shows the reviewer a risk score of 0.87, the recommendation Reject, and a confirm button with the recommendation already selected. Median time on screen is 4 seconds. Audit records this as human oversight. What is the smallest change that makes the oversight real?
Show the full answer Hide the answer
What is being tested
Whether you know the two conditions that make human oversight a control rather than a label. The reviewer must be able to reach a different conclusion from evidence the model has not already summarised, and disagreeing must cost no more than agreeing. This screen fails both.
The mechanism
A score and a pre-selected recommendation do not give the reviewer the claim. They give the reviewer the model's confidence, which a human cannot calibrate: there is no way to know from 0.87 whether the model is usually right at 0.87. So the only judgement available is whether to trust the system, and the deference of people to automated output is one of the better-documented findings in human factors research. A pre-selected default turns that deference into a single click.
The fix is about what is on screen and what the default is, because both are the control. Put the claim's actual fields next to the score (history, amount, the two features that drove the score, the policy clause at issue), and select nothing by default so a decision requires a choice. Then the oversight step has something to oversee.
Why the other options fail
- Typed justification on agreement. Asymmetric friction aimed at the wrong side. It taxes the outcome you may well want and produces a column of free text nobody reads, and within a week the text is the same sentence pasted 400 times. Make disagreement cheap instead of making agreement expensive.
- More reviewers until time on screen rises. Time on screen is a symptom. A reviewer given 30 seconds and nothing but a score still has nothing to assess, so you have bought the same non-control at higher cost. Capacity matters, and it is the next problem, not this one.
- A second confirming reviewer. Two people looking at the same screen with the same pre-selected default produce correlated confirmations. You have doubled the cost of the step and added no independence, which is the degenerate form of four-eyes approval.
When this is the wrong answer
When the human's real job is triage rather than judgement. If 40000 cases a day arrive and ten people are available, no screen design makes per-case oversight possible, and the honest design is explicit sampling: a stated percentage reviewed deeply, chosen by a rule that includes a random component, with the rest decided automatically. That is a legitimate design. What is not legitimate is calling the unsampled remainder overseen.
The monitoring that keeps this honest in production is a floor, not a ceiling: alert when the disagreement rate falls below a threshold such as 2% over a week, because a review step that never disagrees is either unnecessary or not happening.