A lending team's fairness dashboard shows approval-rate parity across protected groups, refreshed monthly, green for six months. Review it. What would you remove, what would you change, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
A lender needs to answer three questions, and the dashboard answers one: are outcomes distributed acceptably, is the model using something it should not, and can an adverse decision be explained to the person it affected? Only the first is on the screen.
What I would change first
One fairness metric is a choice, not a summary, and approval-rate parity is a particular choice with a known consequence: it can be satisfied while the approved populations differ systematically in the terms they receive. Two applicants approved is parity; one at 6% and one at 19% is not equal treatment.
So: report parity alongside error-rate differences — false-positive and false-negative rates by group — and outcome quality among the approved, meaning rate, limit and default rate. The formal result that matters here is that these criteria cannot generally all be satisfied at once when base rates differ, so choosing among them is an explicit, documented decision rather than a technical default. A dashboard that hides the choice hides the decision.
What I would remove
Monthly refresh as the control. A month of lending decisions is a month of people affected, and for smaller groups a month is too few decisions to be significant anyway — so the cadence is simultaneously too slow to act on and too fast to be meaningful. Replace it with continuous computation and alerting on a drift threshold, with confidence intervals shown, so a green cell on 40 decisions is visibly different from a green cell on 40,000.
What I would add
- Proxy analysis. The model may not use a protected attribute and may use postcode, device type or employer, which carry it. The check is whether the attribute is predictable from the features, and it is the omission that most often turns a green dashboard into a finding.
- Reason codes on adverse decisions, because explainability is a separate obligation from fairness and is frequently assumed to be covered by it.
- The pre-model and post-model steps: the marketing that determined who applied, and the manual override that staff apply afterwards. Overrides are where measured fairness is routinely undone, and they are almost never in scope.
What I would leave alone
The dashboard's existence and its ownership. A team monitoring something is in a better position than a team monitoring nothing, and the review should not read as a reason to stop. The green history is also genuine evidence for the one metric it covers, and should be kept as such.
When this is enough as it stands
For a low-stakes, reversible decision with no regulatory duty — which items to show, which email to send — a single parity metric with a threshold is proportionate, and the apparatus above would be governance for its own sake. The line is whether an individual can be materially harmed by one decision and cannot easily reverse it. Lending is over that line. That is the test to apply, not the sophistication of the model.
What the fuller apparatus costs is worth stating rather than assuming: proxy analysis and outcome-quality reporting need the protected attribute to be collected and retained, which is itself a privacy decision some jurisdictions constrain, and the extra metrics will conflict with each other by construction. The failure that follows is a dashboard nobody can act on, where four indicators disagree and no one has been given the authority to choose between them. Choose the criterion in advance and document why — fair lending supervision has expected exactly that written rationale since well before 2020, and it is the artefact most often missing.