Fairness & Bias advanced 9 min read 6 flashcards

Fairness Dynamics and Delayed Impact

Why a constraint that equalises a decision has no built-in relationship to the outcome it is meant to improve, how deployed predictors reshape the distribution they are scored against, and what fairness transfer failures look like in real deployments.

A lender equalises approval rates across two groups. Some years on, average credit scores in the group the intervention was meant to help are lower than they would have been under the unconstrained policy, because the extra loans went to applicants who then defaulted. This is not a bug in the implementation. Static fairness criteria constrain a decision at one moment and say nothing about the trajectory of the people subject to it, and in a one-step model of outcomes, common criteria can produce active harm in regimes where the unconstrained objective would not (Liu et al., 2018, Delayed Impact of Fair Machine Learning, ICML, arXiv:1803.04383).

The outcome curve

Model the effect of a selection rate on a group's average welfare, say mean credit score. Lending to nobody changes nothing. Lending to the most creditworthy raises the average. Keep loosening and you start approving applicants who default, and the curve turns over. Plotted against selection rate, group welfare rises, peaks, and falls, which partitions the space into three regimes: relative improvement, relative harm (worse than the optimum for that group but still better than no intervention), and active harm (worse than doing nothing).

A fairness constraint picks a point on that curve by arithmetic, not by intent. Demographic parity forces whatever selection rate equalises the groups, and whether that lands before or past the peak depends on the group's score distribution, which the constraint never consults. Equal opportunity lands somewhere else, equally blindly. The useful discipline is to ask which regime your constraint puts each group in, which requires an outcome model and an estimate of the decision-to-outcome lag, and is a different exercise from reporting a disparity.

The predictor changes the distribution

Deployment is not a draw from a fixed world. Predictive policing trained on discovered crime, that is arrests, sends patrols back to the neighbourhoods it already patrols, which produces more discovered crime there, which confirms the model. Ensign and colleagues gave this a Pólya-urn model, proved the runaway behaviour follows from the update rule rather than from any bias in the initial data, and showed the correction: reweight each discovered incident by the probability that it would have been discovered, so that intensively patrolled areas do not get credit for their own coverage (Ensign et al., 2018, Runaway Feedback Loops in Predictive Policing, FAT*, arXiv:1706.09847).

The general statement is performative prediction: when predictions inform decisions that influence the outcome, the risk being minimised depends on the model that is deployed, so there is no fixed distribution to be optimal on. The relevant equilibrium is performative stability, a model that is optimal for the distribution it itself induces, reachable under conditions by repeated risk minimisation (Perdomo et al., 2020, Performative Prediction, ICML, arXiv:2002.06673). Regulation has started to notice: the EU AI Act requires that high-risk systems which continue to learn after deployment be built to minimise the influence of biased outputs on future inputs, naming feedback loops directly (EU AI Act, Article 15).

Fairness does not transfer

Even without feedback, a model that satisfies a criterion at one site frequently violates it at the next. In two real medical deployments, the shifts encountered were more structurally complex than the covariate-shift and label-shift idealisations the literature leans on, and fairness properties measured at the source did not survive the move. The practical contribution is diagnostic: conditional independence tests characterise which kind of shift you are facing, and which kind you are facing determines which properties can be expected to transfer at all (Schrouff et al., 2022, Diagnosing Failures of Fairness Transfer Across Distribution Shift in Real-World Medical Settings, NeurIPS, arXiv:2202.01034). Simulation studies of long-run fairness reach the same conclusion from the other direction, and the agent-based tooling for running them is public (D'Amour et al., 2020, Fairness Is Not Static, FAT*; ML-fairness-gym).

What this implies for a deployed system

Measure the outcome, not only the decision. That means instrumenting the thing the intervention was supposed to improve (repayment, completion, recovery) on the horizon it actually moves on, and accepting that the first honest read arrives quarters after launch.

Put fairness metrics in production monitoring with the retraining loop inside the measured system, not outside it, since a loop that retrains on its own decisions is the mechanism above (see slice-based monitoring and alert design). And simulate before shipping where the stakes justify it, with the simulation's assumptions written down as assumptions.

When it breaks

The outcome curve is a model, and its shape is rarely known. The three-regime picture is a conclusion from a parameterised model of how decisions change welfare. Which regime a real policy occupies depends on parameters you are estimating from the same history that produced the disparity.

Simulation inherits the modeller's beliefs. An agent-based study of long-run fairness is an argument about dynamics, not evidence about them, and its conclusions follow from its transition assumptions. Use it to find policies that fail under plausible assumptions rather than to certify one that succeeds.

Long-run claims are unfalsifiable on a quarterly cadence. The honest position is that the long-run effect is unmeasured, with the shortest leading indicator you can find reported alongside it. Teams under pressure to show fairness progress will reach for the decision-time metric because it reports immediately, which is precisely the substitution this concept argues against.

Not every harm is welfare-shaped. Dignitary harms, stigma and the cost of being wrongly flagged do not sit on a credit-score axis, and a dynamic analysis that only tracks a scalar will score an intervention as improvement while the affected group experiences the opposite.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Liu et al., 2018, Delayed Impact of Fair Machine Learning, ICML, arXiv:1803.04383 arxiv.org
  2. Ensign et al., 2018, Runaway Feedback Loops in Predictive Policing, FAT*, arXiv:1706.09847 arxiv.org
  3. Perdomo et al., 2020, Performative Prediction, ICML, arXiv:2002.06673 arxiv.org
  4. EU AI Act, Article 15 artificialintelligenceact.eu
  5. Schrouff et al., 2022, Diagnosing Failures of Fairness Transfer Across Distribution Shift in Real-World Medical Settings, NeurIPS, arXiv:2202.01034 arxiv.org
  6. D'Amour et al., 2020, Fairness Is Not Static, FAT*; ML-fairness-gym github.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track