Multicalibration and Subgroup Guarantees
Why satisfying a fairness constraint on each protected attribute separately leaves intersections unprotected, what it means to ask for calibration on every computationally identifiable subgroup at once, and how a post-processing loop achieves it.
Constrain a risk model to equalise false positive rates across race, then across sex. Both constraints hold. Now look at Black women over 60: the rate there can be far outside anything the marginal constraints permit, because nothing in them refers to the intersection. Worse, a model can be built deliberately to exploit this, satisfying every marginal constraint while concentrating error on a structured subgroup that no reported metric covers. Kearns and colleagues named it fairness gerrymandering (Kearns et al., 2018, Preventing Fairness Gerrymandering: Auditing and Learning for Subgroup Fairness, ICML, arXiv:1711.05144).
The obvious fix, constrain every subgroup, does not survive contact with combinatorics. With \(k\) binary protected attributes there are \(3^k\) subgroups defined by conjunctions, most of them too small to estimate anything on, and the paper proves that auditing for subgroup violations is equivalent to weak agnostic learning over the subgroup class, so it is computationally hard in the worst case even for simple classes.
Calibration on every identifiable subgroup
The response that scales is to fix a class of subgroups and demand calibration within all of them simultaneously. Let \(p: X \to [0,1]\) be a predictor and \(\mathcal{C}\) a collection of subsets of the input space, each identifiable by some bounded computation. Then \(p\) is \((\mathcal{C}, \alpha)\)-multicalibrated if for every \(S \in \mathcal{C}\) that is not vanishingly rare and every score value \(v\),
Read it carefully, because it is stronger than it looks. It is not "the average prediction is right for each group". It is "at every score level, within every group in the class, the score means what it says" (Hébert-Johnson et al., 2018, Multicalibration: Calibration for the (Computationally-Identifiable) Masses, ICML, arXiv:1711.08513). A score of 0.3 has to correspond to a 30 percent base rate among the people in \(S\) who received a 0.3, not just on average over \(S\).
The class \(\mathcal{C}\) carries the whole value judgement. It can contain conjunctions of protected attributes, or every subgroup a decision tree of depth 4 can pick out, or every group a linear threshold on the features can describe. Richer classes give stronger guarantees and need more data, and the complexity of learning a multicalibrated predictor turns out to match the complexity of weak agnostic learning for that class, the same quantity that makes auditing hard.
Getting there by patching
The algorithm is a loop, and its shape explains the guarantee. Audit the current predictor for a subgroup and score bucket where calibration is violated by more than \(\alpha\). If one exists, shift the predictions in that bucket toward the observed outcome rate in that subgroup. Re-audit. Each successful patch reduces a squared-error potential by an amount bounded below, and squared error cannot fall below zero, so the number of rounds is bounded independently of how many subgroups there are. The cost of the guarantee is paid in audit queries, not in rounds.
Multiaccuracy is the weaker sibling: match conditional means per subgroup rather than full calibration. It needs only black-box access to the predictor and a modest labelled audit set, which makes it the version you can actually run against a vendor model you cannot retrain (Kim, Ghorbani and Zou, 2019, Multiaccuracy: Black-Box Post-Processing for Fairness in Classification, AIES, arXiv:1805.12317).
What it does not buy
Multicalibration is a member of the sufficiency family, so it does not escape the impossibility result (see group fairness criteria and why they conflict). A multicalibrated score will not in general equalise error rates across groups, and if your harm argument is about false positives falling on one group, multicalibration is not the criterion that addresses it.
The absolute error bound also hides a distributional problem. A tolerance of \(\alpha = 0.01\) is mild for a group with a 50 percent base rate and severe for one with a 2 percent base rate, where it permits a 50 percent relative error. Proportional multicalibration reformulates the constraint in relative terms for exactly this reason, and was motivated by clinical risk scores where low-prevalence groups are the ones under discussion (La Cava, Lett and Wan, 2023, Fair Admission Risk Prediction with Proportional Multicalibration, CHIL, PMLR 209).
When it breaks
Choosing \(\mathcal{C}\) is choosing who is protected. The formalism makes the choice explicit, which is an improvement, and it does not make the choice for you. A class that cannot express the subgroup someone is worried about gives them no guarantee at all.
The audit set is the binding constraint in practice. Guarantees hold for subgroups with enough mass; a subgroup that is 0.3 percent of your labelled data gets a vacuous bound. Rich-subgroup methods trained with heuristics in place of learning oracles do converge quickly and buy real fairness for mild accuracy cost, but the empirical work also shows that optimising accuracy under marginal constraints alone leaves substantial subgroup unfairness behind (Kearns et al., 2019, An Empirical Study of Rich Subgroup Fairness for Machine Learning, FAT*, arXiv:1808.08166).
Patching is post-processing, with post-processing's legal problem. If the patch is keyed on a protected attribute at inference time, it may be disparate treatment regardless of its effect (see disparate impact and the legal frame). Subgroups defined by non-protected proxies avoid the issue and weaken the guarantee.
A calibrated score for the wrong target is still the wrong target. Multicalibration conditions on the labels you have. If the label is health spending and the construct is health need, every subgroup gets a faithful score for the wrong quantity.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Kearns et al., 2018, Preventing Fairness Gerrymandering: Auditing and Learning for Subgroup Fairness, ICML, arXiv:1711.05144 arxiv.org
- Hébert-Johnson et al., 2018, Multicalibration: Calibration for the (Computationally-Identifiable) Masses, ICML, arXiv:1711.08513 arxiv.org
- Kim, Ghorbani and Zou, 2019, Multiaccuracy: Black-Box Post-Processing for Fairness in Classification, AIES, arXiv:1805.12317 arxiv.org
- La Cava, Lett and Wan, 2023, Fair Admission Risk Prediction with Proportional Multicalibration, CHIL, PMLR 209 proceedings.mlr.press
- Kearns et al., 2019, An Empirical Study of Rich Subgroup Fairness for Machine Learning, FAT*, arXiv:1808.08166 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.