Instrumental Variables
Using a source of variation that affects treatment but has no other path to the outcome, which recovers a causal effect despite unmeasured confounding, for a subpopulation you cannot identify.
Every method that adjusts for covariates assumes you measured the confounders. An instrument makes a different bet: find a source of variation in treatment that is as good as random, and use only that variation. Unmeasured confounding is then irrelevant to the part of treatment you exploit.
\(Z\) is a valid instrument for the effect of \(T\) on \(Y\) when three conditions hold. Relevance: \(Z\) affects \(T\), which is testable. Exclusion: \(Z\) affects \(Y\) only through \(T\), which is not testable. Independence: \(Z\) is unconfounded with \(Y\)'s unobserved causes, also not testable. Two of the three assumptions are assertions defended by argument.
The estimator, and what it is dividing
With a binary instrument, the Wald estimator is
The numerator is the intention-to-treat effect: how much the instrument moved the outcome. The denominator is how much the instrument moved treatment. Dividing rescales the ITT to a per-unit-of-treatment effect. If an encouragement raised adoption by 20 percentage points and raised the outcome by 2 points, the implied effect of adoption is 10 points.
Two-stage least squares generalises this: regress \(T\) on \(Z\) and covariates, then regress \(Y\) on the fitted \(\hat{T}\). The fitted values contain only the variation in \(T\) explained by the instrument, which is by assumption unconfounded. The confounded part of \(T\) is discarded, along with most of the precision.
LATE: whose effect is it
Under a monotonicity assumption, that no unit is a defier who does the opposite of what the instrument encourages, IV identifies the local average treatment effect: the effect among compliers, the units whose treatment status was actually changed by the instrument (Imbens and Angrist, 1994, Econometrica 62(2)).
This is the most under-reported feature of the method. Always-takers and never-takers contribute nothing, and compliers are not identifiable individually, so the estimand is an average over a subpopulation you cannot describe or count directly. A judge-assignment instrument estimates the effect on defendants whose sentence depended on which judge they drew. A distance-to-hospital instrument estimates the effect on patients whose treatment depended on travel time. Different instruments for the same treatment recover genuinely different quantities, and disagreement between two IV estimates is not necessarily a contradiction.
When it breaks
Weak instruments are worse than no instrument. When the first stage is weak, the denominator is near zero and the estimator becomes badly biased toward OLS, with confidence intervals that do not have their stated coverage. The conventional screen is a first-stage F-statistic above 10, and recent work argues that threshold is too permissive for the coverage it claims. A weak-instrument IV analysis is not a noisy answer, it is a confidently wrong one.
Exclusion is where the argument lives and where it usually fails. Any path from \(Z\) to \(Y\) that does not pass through \(T\) invalidates the design. Distance to a hospital is correlated with income, which affects health. Weather as an instrument for outdoor activity affects mood directly. The assumption cannot be tested with one instrument; with more instruments than endogenous regressors, over-identification tests give a weak check that is often misread as validation.
Small violations are amplified. Because IV divides by the first-stage strength, a small direct effect of \(Z\) on \(Y\) is scaled by the same factor as the treatment effect. With a first stage of 0.1, a bias of 0.01 in the numerator becomes 0.1 in the estimate. Weak instruments and exclusion violations compound each other multiplicatively.
Randomised encouragement is the clean case, and it is available more often than people assume. Randomising an invitation, a prompt, or a default gives an instrument that is valid by construction, and it is the standard design when the treatment itself cannot be randomised. It converts an untestable assumption into a design property, which is the strongest move available in this area.
7 flashcards for this concept
Click a card to reveal the answer.