Search Evaluation
Pooling and judgments, nDCG and MRR, interleaving, online metrics, and why offline gains vanish online.
4concepts
54flashcards
28minutes of reading
- 01 Test Collections, Pooling and Judgment Bias The Cranfield paradigm made retrieval a measurable science, and the pooling shortcut that makes it affordable quietly penalises any system unlike the ones that built the pool.
- 02 nDCG, MRR and Graded Relevance The main ranking metrics differ in what they assume about the user, and choosing one is choosing a model of how far someone reads and what they are looking for.