Interference and Network Effects in Experiments
When one unit's treatment affects another unit's outcome, individual randomisation measures a quantity that is neither the treatment effect nor zero, and the standard designs trade bias against a large loss of power.
A marketplace tests a ranking change that surfaces certain listings more prominently. Treated buyers see the new ranking and convert 8% more often. The experiment reports an 8% lift. Rolled out to everyone, the lift is 1%, because the treated buyers were partly taking supply and attention from control buyers rather than creating new demand. The experiment measured a redistribution and reported it as growth.
This is interference, the violation of the "no interference" half of SUTVA. It is not an edge case in any system where units share a finite resource or influence each other: marketplaces, social products, ad auctions, delivery networks, and anything with a shared budget or inventory.
Why the individually randomised estimate is not simply biased
Under interference, unit \(i\)'s outcome depends on the entire assignment vector, so \(Y_i(1)\) and \(Y_i(0)\) are not well defined without specifying what everyone else received. The individually randomised comparison estimates the difference between "treated in a world that is 50% treated" and "untreated in a world that is 50% treated". Neither is the quantity a launch decision needs, which is "everyone treated" versus "nobody treated".
The direction is predictable from the mechanism. Under competition for a shared resource, the treated arm gains partly at the control arm's expense, so the estimate overstates the global effect: the control group is depressed, widening the gap. Under positive spillover, treated users influence their untreated friends, so the control arm is lifted and the estimate understates the global effect.
The standard designs, and what each costs
Cluster randomisation. Randomise groups within which interference is concentrated: geographic regions for a marketplace, social communities detected on the graph, whole accounts for a B2B product. Interference within a cluster is absorbed into the cluster's outcome; interference across clusters is assumed away. The cost is severe and often underestimated: the effective sample size becomes the number of clusters, not users, and the design effect \(1 + (m-1)\rho\) can inflate the required sample by an order of magnitude. Fifty regions is fifty units, whatever the user count.
Switchback (time-based) randomisation. Alternate the treatment for the entire system across time intervals. Everyone experiences the same condition simultaneously, so within-period interference is fully contained, which suits marketplaces where the interference is through instantaneous supply and demand. The costs are carryover between adjacent periods, requiring burn-in windows, and confounding with time-of-day and day-of-week, requiring balanced designs and a long run.
Ego-cluster and graph-aware designs. Randomise a focal user together with their neighbourhood, so a treated ego sits in a mostly treated environment. This gives a cleaner exposure contrast on social graphs than naive assignment while using far fewer units than full community clustering.
Two-stage designs. Randomise clusters to a treatment saturation level, then randomise units within them. Varying saturation across clusters allows the spillover effect to be estimated rather than only avoided, which is the only way to learn the shape of the interference rather than to design around it.
When it breaks
Interference is usually invisible in the metrics. The individually randomised experiment reports a clean, significant, well-powered result. Nothing in the analysis indicates that the estimand is wrong. Detection requires either a design that could reveal it, such as varying saturation, or a prior structural argument about resource sharing.
Clusters leak. Geographic clusters share users who travel; social clusters share bridging ties. The assumption is not that interference is zero across clusters but that it is small, and that claim needs an argument.
The power loss makes clustering unaffordable for small effects. With 60 regions, an experiment can detect only large effects, so the honest choice is often between a biased estimate of a small effect and no estimate at all. Stating that trade explicitly is better than defaulting to the individually randomised design because it fits the power budget.
Switchbacks assume the effect acts fast. A change whose impact accumulates over days, such as anything affecting retention or reputation, cannot be measured in hour-long alternating periods, because the carryover never clears.
6 flashcards for this concept
Click a card to reveal the answer.