Search Evaluation
4concepts
27flashcards
28minutes of reading
- 01 Test Collections, Pooling and Judgment Bias The Cranfield paradigm made retrieval a measurable science, and the pooling shortcut that makes it affordable quietly penalises any system unlike the ones that built the pool.
- 02 nDCG, MRR and Graded Relevance The main ranking metrics differ in what they assume about the user, and choosing one is choosing a model of how far someone reads and what they are looking for.
- 03 Interleaving and Online Evaluation Mixing two rankers' results into a single list and attributing clicks gives a within-user paired comparison that detects differences far faster than an A/B test on the same traffic.
- 04 Why Offline Gains Vanish Online The recurring experience that an offline nDCG improvement produces no measurable online effect, and the four distinct mechanisms that cause it.