Disparate Impact and the Legal Frame
How US anti-discrimination doctrine actually assigns liability, why the four-fifths rule is a screening heuristic rather than a standard, and the places where a passing fairness metric and a successful legal defence come apart.
A screening model ships with an impact ratio of 0.83 across race categories, above the familiar 0.8 line, and the team records the number as evidence of compliance. The number is worth having. It is not a defence, it is not a test the law recognises as dispositive, and in the one jurisdiction that mandates publishing it, nearly every audit that got published cleared it (Wright et al., 2024, Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability, FAccT). Fairness metrics and legal exposure are related quantities, and a team that treats them as the same quantity will be surprised in both directions.
Two doctrines that behave differently
Disparate treatment is differential handling because of a protected characteristic. Intent matters, and a practice that uses the attribute in the decision is exposed regardless of its effect, which is why group-specific thresholds are usually unavailable even when they are the technically cleanest correction (see mitigation at pre-, in- and post-processing).
Disparate impact is a facially neutral practice with a disproportionate effect on a protected group. Griggs v. Duke Power Co., 401 U.S. 424 (1971) established it under Title VII: an aptitude test and a high-school diploma requirement screened out Black applicants at much higher rates and could not be shown to predict job performance. No intent was required, and none was found.
The structure of an impact claim is a three-step burden shift. The plaintiff shows the disparity. The employer shows the practice is job-related and consistent with business necessity. The plaintiff then shows a less discriminatory alternative that serves the same interest. That last step is where machine learning sits awkwardly: a model is one point in a large space of models, so "a less discriminatory alternative exists" is often provable by just training one, and the doctrine never contemplated a defendant who could enumerate the alternatives (Barocas and Selbst, 2016, Big Data's Disparate Impact, 104 California Law Review 671).
What the four-fifths rule actually says
The rule lives in the EEOC's selection guidelines: a selection rate for any race, sex or ethnic group below four-fifths of the rate of the highest group "will generally be regarded by the Federal enforcement agencies as evidence of adverse impact" (29 CFR 1607.4(D)). Selection rates of 0.100 and 0.125 give a ratio of exactly 0.80.
The regulation then qualifies itself twice. Smaller differences may still constitute adverse impact where they are significant in both statistical and practical terms. Larger differences may not, where they rest on small numbers and are not statistically significant. So the threshold is a flag that invites scrutiny, not a line that settles anything, and reading it as a pass mark inverts its function. It is a ratio of point estimates with no sample size in it, which makes it loud on small applicant pools and quiet on large ones, exactly backwards from what a test of evidence should do (Watkins, McKenna and Chen, 2022, The four-fifths rule is not disparate impact, arXiv:2202.09519; Kim and Raghavan, 2023, Limitations of the "Four-Fifths Rule" and Statistical Parity Tests for Measuring Fairness, NeurIPS).
Who is liable, and who has to measure
Two developments changed the practical picture for anyone shipping a model that screens people.
Vendors are no longer obviously outside the frame. In Mobley v. Workday the Northern District of California let claims proceed on the theory that a vendor whose tool performs a traditional hiring function is acting as an agent of its employer customers, and in May 2025 the court granted preliminary certification of an ADEA collective reaching screening decisions back to 2020. The theory makes the party that built the model a defendant, not just the party that bought it.
Measurement is being mandated. New York City's Local Law 144 requires an annual independent bias audit of automated employment decision tools, computing selection rates and impact ratios by sex and race or ethnicity, with the results posted publicly; enforcement began 5 July 2023. The compliance study above found 18 posted audit reports and 13 transparency notices across 391 employers checked, and because employers decide whether their own tool is in scope, a missing report cannot be read as non-compliance. In the EU, Article 10(5) of the AI Act runs the other way, creating a narrow legal basis to process special-category data precisely so that bias can be detected and corrected, under conditions including that the work cannot be done with anonymised or synthetic data and that the data is deleted afterwards (EU AI Act, Article 10).
When it breaks
The denominator is a decision, not a fact. A selection rate needs a pool, and who counts as an applicant (everyone who clicked, everyone who completed, everyone the sourcing tool surfaced) moves the ratio more than most modelling choices do. Fix the definition before the number, and record it.
Positive discrimination is still discrimination. A model that favours a protected group is a disparate treatment problem, not a fairness win, and an audit that reports only the worst-off group will miss it.
Clearing the ratio on the model does not clear the system. Liability attaches to the practice. A compliant scoring stage feeding a recruiter who overrides it in a patterned way produces a disparate outcome from a fair model, and the discovery process will look at the outcome.
This is one jurisdiction's doctrine. EU and UK equality law, and sector rules in lending, insurance and housing, have different structures and different available defences. Treat the four-fifths arithmetic as a US employment artefact that other regimes do not share.
The metric is the beginning of the argument. What survives scrutiny is the record: why this target variable, why these features, which alternatives were tried and what they cost. Teams that keep that record have a business-necessity story. Teams with a dashboard have a number.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Wright et al., 2024, Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability, FAccT dl.acm.org
- Barocas and Selbst, 2016, Big Data's Disparate Impact, 104 California Law Review 671 papers.ssrn.com
- 29 CFR 1607.4(D) ecfr.gov
- Watkins, McKenna and Chen, 2022, The four-fifths rule is not disparate impact, arXiv:2202.09519 arxiv.org
- Kim and Raghavan, 2023, Limitations of the "Four-Fifths Rule" and Statistical Parity Tests for Measuring Fairness, NeurIPS openreview.net
- EU AI Act, Article 10 artificialintelligenceact.eu
6 flashcards for this concept
Click a card to reveal the answer.