Reasoning & Evaluation
25 min
Error Bars for Evals: Why Most Benchmark Differences Are Noise
A 250-question benchmark carries a standard error of about three percentage points. Most of the model comparisons published on top of such benchmarks cannot distinguish the models they are comparing. Evaluations are experiments, …