AI Diffusion & Labour advanced 7 min read 12 flashcards

Measuring Diffusion Honestly

Why most claims about AI's economic impact rest on evidence that cannot support them, the hierarchy of evidence quality, and the specific questions to ask of any figure.

Claims about AI's effect on work circulate with a confidence the evidence does not support, in both directions. The useful skill is not having a view but knowing what kind of evidence a given claim rests on and what that kind of evidence can establish.

The hierarchy

Randomised controlled trials are the strongest available and are narrow: a specific task, a specific population, a specific tool version, over a short period. They establish causation within that scope, and generalising beyond it is an assumption. The METR and Noy-Zhang studies are in this class and reach different conclusions because their scopes differ.

Quasi-experimental field studies exploit staggered rollouts to estimate effects in real workplaces, which is more realistic and less controlled. The customer support result is the best-known example.

Usage telemetry shows where systems are actually applied, at scale, without self-report bias. It cannot establish effect, only application.

Aggregate economic statistics are the outcome anyone actually cares about and are far too slow and too noisy to attribute to a single technology at this stage. Any claim that AI has or has not moved national productivity is currently outrunning the data.

Surveys and self-report are the weakest, and are the basis of most published claims. The METR finding, that participants believed they were 20 percent faster while being 19 percent slower, is direct evidence that self-report on this specific question is unreliable in a measurable direction.

Vendor case studies sit below that, since they are selected for a favourable result by a party with an interest in it.

Questions to ask of a figure

What was measured, and is it the outcome or a proxy? Compared against what baseline? Over what period, and does it include learning and novelty effects? On which population, and would it hold for a different skill level? Who conducted it and what is their interest? Is the effect on the individual, the team, or the firm, and does the claim match the level at which it was measured?

A figure that survives those questions is worth carrying. Most do not, and the failure is usually at the baseline or the level.

When it breaks

Both optimism and pessimism outrun the evidence. Claims of imminent mass displacement and claims that nothing will change are equally unsupported by the current data, and the honest position is that the effects are real, measurable at task level, and not yet resolvable at the aggregate level.

Historical analogies are used as arguments. Previous general-purpose technologies did display long adoption lags followed by measurable effects, which is a reason to expect delay rather than a reason to expect any particular outcome. The analogy constrains the timeline more than the destination.

The technology is moving during measurement. A study of a 2024 model says something about that model, and capability moves fast enough that a result can be stale before publication. This is a genuine methodological problem for the field and not a criticism of any individual study.

Distributional effects are underreported relative to aggregate ones. Aggregate employment stability is the easy headline and is compatible with substantial disruption for specific groups. Asking who, not how many, is what makes the evidence useful for anyone deciding anything.

Check yourself

12 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track