concept

Fairness Definition Choice

also called Incompatible Fairness Metrics, Which Fairness

The unavoidable selection between mathematically incompatible fairness definitions - a legal and business determination that a team makes implicitly if it does not make it explicitly.

digitackofairnesssubgroupspolicy

Fairness is not one property. Equal outcome rates across groups, equal error rates, equal treatment of similar individuals, and calibration within groups are different definitions — and they are mathematically incompatible except in special cases.

Satisfying one means failing another, which means every fairness assessment has already chosen, whether or not anyone said so.

Why it matters

A team that picks a definition implicitly has made a policy decision without saying so, and it will be examined later by someone applying a different one. The choice is a legal and business determination that depends on the domain, the jurisdiction and the harm being guarded against.

Making it explicit before the evaluation is designed is what makes the result defensible.

Implementation patterns

  • State the definition applied, with its rationale, before the evaluation is built.
  • Report against more than one where they conflict, so the trade-off is visible rather than hidden by the choice.
  • Subgroup evaluation as a required gate, since aggregate performance can be excellent while a subgroup is served badly, and that subgroup is frequently the one that matters legally.
  • A governed dataset for assessment, separately controlled, since measuring fairness across a protected characteristic requires knowing it and collecting it may be restricted — a tension that must be designed for rather than improvised.
  • Proxy detection, since a model can reproduce a protected characteristic through correlated features without ever seeing it. Excluding the characteristic does not remove the effect, which is the most common misunderstanding.
  • Post-deployment monitoring, since the population changes and a fair-at-launch model drifts.
  • An appeal path reaching a human with authority to change the outcome, since an appeal returning the same automated answer is not an appeal.

Industry example

Insurers such as Digit and Acko price on risk, which is both legitimate and regulated. The line between a legitimate risk factor and a proxy for a protected characteristic is a regulatory judgement that varies by jurisdiction and changes over time — which makes the factor set a governed artefact with an effective date, reviewable and defensible, rather than whatever the model found predictive.

Failure scenarios

  • A definition chosen implicitly, by whichever metric was easiest to compute.
  • Aggregate-only evaluation, hiding subgroup harm.
  • The protected characteristic excluded and reproduced through proxies.
  • No assessment dataset, making measurement impossible.
  • The factor set driven purely by predictive power, with no regulatory review.

Trade-offs

Reporting against multiple definitions makes the trade-off visible, which is honest and commercially uncomfortable — it documents that the system is unfair by some measure, because it must be.

The alternative is a single reported metric that conceals the choice, which is more comfortable and less defensible. The documentation that carries the weight is what was tested, against which definition, on which subgroups, with what result, and what was accepted — because its absence means the answer to a challenge is a re-analysis conducted under pressure.

Interview question

"Your model has equal accuracy across two groups and different approval rates. Is it fair? Tell me who decides, and what you would have needed to record before this question was asked."