Observational Causal Methods intermediate 8 min read 7 flashcards

Sensitivity Analysis and E-Values

Asking how strong an unmeasured confounder would have to be to explain away an observational result, and the three standard ways to quantify that, from Rosenbaum's gamma to the E-value to partial R-squared bounds.

Users who adopted the team workspace feature were 1.8 times as likely to renew, after adjusting for plan, seat count, tenure and usage. The obvious objection is that teams who adopt collaboration features were already more committed. No adjustment for measured covariates can rule that out. What an analysis can do is put a number on the objection: how strong would that commitment have to be, in its relation to both adoption and renewal, to produce a 1.8 from a true 1.0?

That reframing goes back to Cornfield's argument about smoking and lung cancer, and it turns an unanswerable "is there unmeasured confounding?" into an answerable "how much would matter?". Three families of methods dominate.

Rosenbaum's gamma for matched designs

In a matched-pair study, randomisation would give each unit in a pair a treatment probability of \(1/2\). Rosenbaum's model allows hidden bias: two units with identical observed covariates may have odds of treatment differing by a factor of at most \(\Gamma \ge 1\), so within a pair

\[\frac{1}{1 + \Gamma} \;\le\; \pi_i \;\le\; \frac{\Gamma}{1 + \Gamma}\]

(Rosenbaum, 1987, Sensitivity Analysis for Certain Permutation Inferences in Matched Observational Studies, Biometrika 74(1)). At \(\Gamma = 2\) each unit's treatment probability can be anywhere between \(1/3\) and \(2/3\). For each \(\Gamma\) the method computes the worst-case p-value, and the report is the largest \(\Gamma\) at which the result remains significant. A study that loses significance at \(\Gamma = 1.1\) is fragile; one that survives \(\Gamma = 5\) requires a hidden confounder that makes treatment five times likelier for otherwise identical units.

The E-value

VanderWeele and Ding proposed a single summary that needs no model of the confounder (VanderWeele and Ding, 2017, Sensitivity Analysis in Observational Research: Introducing the E-Value, Annals of Internal Medicine 167(4)). For an observed risk ratio \(RR > 1\),

\[E = RR + \sqrt{RR\,(RR - 1)}\]

(for \(RR < 1\), apply the formula to \(1/RR\)). It rests on a bound: an unmeasured confounder \(U\) with risk ratio \(RR_{EU}\) relating treatment to \(U\) and \(RR_{UD}\) relating \(U\) to the outcome can shift the observed risk ratio by at most

\[B = \frac{RR_{EU}\cdot RR_{UD}}{RR_{EU} + RR_{UD} - 1}\]

(Ding and VanderWeele, 2016, Sensitivity Analysis Without Assumptions, Epidemiology 27(3)). The E-value is the common value of the two risk ratios at which \(B\) equals the observed effect.

For the workspace result, \(E = 1.8 + \sqrt{1.8 \times 0.8} = 1.8 + 1.2 = 3.0\). A confounder associated with both adoption and renewal by a risk ratio of 3, above and beyond the measured covariates, could fully explain the 1.8; a weaker one could not. Check: \(B = 9/5 = 1.8\). The bound also works asymmetrically: a confounder with \(RR_{UD} = 6\) needs only \(RR_{EU} \approx 2.14\). If the lower confidence limit is 1.3, its E-value is \(1.3 + \sqrt{0.39} \approx 1.92\), the strength needed to make the interval include 1. Report both.

Omitted variable bias in partial R-squared

For linear regression, Cinelli and Hazlett re-express the classic omitted-variable bias formula in terms of how much residual variance a confounder would explain in the treatment and in the outcome, and summarise it with a robustness value (Cinelli and Hazlett, 2020, Making Sense of Sensitivity: Extending Omitted Variable Bias, JRSS-B 82(1)). Their key practical move is benchmarking: bounding the confounder's strength as a multiple of an observed covariate's. "A confounder three times as strong as tenure would be needed" is a claim a domain expert can argue with; "a partial \(R^2\) of 0.07" is not.

When it breaks

E-values are easy to misread. Ioannidis, Tan and Blum argued that because the E-value is a monotone transform of the effect estimate it adds no new information, and that it invites readers to declare results robust without asking whether confounders of that strength are plausible (Ioannidis, Tan and Blum, 2019, Annals of Internal Medicine 170(2)). VanderWeele, Mathur and Ding replied that the transform is the point: it translates an estimate into the scale on which confounding is argued (2019, Correcting Misinterpretations of the E-Value, Annals of Internal Medicine 170(2)). Either way, an E-value reported without a discussion of plausible confounders is decoration.

No threshold is universal. An E-value of 3 is large in a domain where known confounders have risk ratios near 1.2 and small where selection effects of 5 are routine. Benchmark against measured covariates.

Sensitivity analysis addresses confounding, not other biases. Selection into the sample, measurement error in the outcome and model misspecification each need their own analysis. A robust E-value says nothing about a collider you conditioned on.

Worst-case bounds are pessimistic by design. Rosenbaum's \(\Gamma\) and the E-value both assume the confounder acts in the most damaging direction. Real confounders rarely align perfectly, so these are conservative summaries, not estimates of the actual bias.

Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track