Simpson's Paradox and Choosing an Adjustment Set
The same data can show an effect in every subgroup and the opposite effect in aggregate, and the arithmetic cannot tell you which is right; only the causal structure can.
Treatment A beats treatment B among patients with small stones. It also beats B among patients with large stones. Pooled across both, B beats A. Nothing is miscalculated; both statements are arithmetically correct, and the reversal is a real feature of the numbers.
The reversal happens because the subgroups have different sizes and different base rates, and treatment assignment correlates with subgroup. If A was preferentially given to the harder cases, its pooled performance carries that disadvantage. Aggregation mixes a within-group comparison with a between-group difference in case mix.
The paradox is not a statistical question
The data alone cannot say which number is the right answer. Both aggregations are valid summaries of the same table. Deciding which one answers the causal question requires knowing the role of the grouping variable, and that comes from the causal structure rather than from the numbers.
If the grouping variable is a confounder, a common cause of both treatment and outcome, the disaggregated result is the causal one. Stone size affects which treatment a doctor chooses and also affects recovery, so it sits on a backdoor path and must be blocked.
If the grouping variable is a mediator, caused by treatment and in turn affecting the outcome, the aggregated result is the causal one for the total effect. Conditioning on it removes part of the effect being measured. A drug that works by lowering blood pressure, analysed within levels of blood pressure, appears to do nothing.
The two cases are indistinguishable from the table. They are distinguished by knowing what causes what, which is exactly the content a DAG makes explicit.
Choosing the set, in practice
The workable procedure is short and rarely followed. Draw the graph, including variables you cannot measure. Apply the backdoor criterion to find sets that block all backdoor paths without conditioning on descendants of treatment. Where several valid sets exist, prefer the one that is precisely measured and strongly predictive of the outcome, since that reduces variance without changing identification.
Two heuristics survive contact with reality. Adjust for pre-treatment variables that plausibly cause both treatment and outcome. Do not adjust for anything realised after treatment, in an experiment or outside one, unless mediation is explicitly the question. Most bad controls are caught by the second rule alone.
When it breaks
Continuous versions hide the reversal. The paradox is usually taught with a binary grouping variable and a small table, which makes it visible. With a continuous covariate and a regression, the same phenomenon appears as a coefficient that changes sign when a control is added, and it is frequently interpreted as "the effect was confounded" without asking whether the added control was a mediator or a collider.
Aggregating over time is the same trap. Comparing this quarter to last quarter while the customer mix shifts is Simpson's paradox with time as the grouping variable, and a metric can decline in every segment while rising overall. This is why segment-level reporting is not optional for any metric whose population composition moves.
Neither answer may be the one you want. If the grouping variable is a collider, both the pooled and the stratified estimates are biased and the correct analysis conditions on neither. The choice is not always between two candidate answers.
More granular is not more correct. Stratifying further eventually produces cells with too few observations, where the within-cell estimates are noise. There is a real bias-variance trade in how finely to condition, and taking the most disaggregated view on principle is its own error.
6 flashcards for this concept
Click a card to reveal the answer.