Reasoning & Evaluation
20 min
When the Judge Is Also a Player: LLM-as-Judge, Contamination, and Why Leaderboards Drift
A strong model grading other models looks like a free lunch for evaluation. It is not. Position, verbosity, and self-preference biases plus quietly leaked test sets mean a leaderboard number can move several points without any mo…