Reasoning & Evaluation
24 min
The Ladder of Causation: Why No Amount of Observational Data Climbs It Alone
In 2023 GPT-4 scored 97% on a classic cause-and-effect benchmark and 62% on one that hands it the causal graph and asks it to compute. Both results fit a theorem proved in 2020: data from one rung of Pearl's ladder almost never d…