Reasoning & Evaluation
24 min
The Leaderboard Is Not Your Corpus: Why Top-Ranked Embedding Models Disappoint in Production
Embedding models are chosen from a leaderboard more often than from an experiment, and the leaderboard now publishes training splits for its own test sets. Between contamination, task-family averaging and geometry no benchmark me…