Offline Retrieval Evaluation Without Evaluation Metrics
Fernando Diaz, Andres Ferraro
Abstract
Offline evaluation of information retrieval and recommendation has traditionally focused on distilling the quality of a ranking into a scalar metric such as average precision or normalized discounted cumulative gain. We can use this metric to compare the performance of multiple systems for the same request. Although evaluation metrics provide a convenient summary of system performance, they also collapse subtle differences across users into a single number and can carry assumptions about user behavior and utility not supported across retrieval scenarios. We propose recall-paired preference (RPP), a metric-free evaluation method based on directly computing a preference between ranked lists. RPP simulates multiple user subpopulations per query and compares systems across these pseudo-populations. Our results across multiple search and recommendation tasks demonstrate that RPP substantially improves discriminative power while correlating well with existing metrics and being equally robust to incomplete data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00513942-e2ea-464e-9fcf-e3d1929ffb7dBuilds on1
Related papers
- Evaluation Measures Based on Preference GraphsCharles L. A. Clarke, Chengxi Luo, Mark D. SmuckerSIGIR 2021 · 5 citations
- Good Evaluation Measures based on Document PreferencesTetsuya Sakai, Zhaohao ZengSIGIR 2020 · 14 citations
- Offline Evaluation of Ranked Lists using Parametric Estimation of PropensitiesVishwa Vinay, Manoj Kilaru, David ArbourSIGIR 2022
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 21 citations
- Agreement and Disagreement between True and False-Positive Metrics in Recommender Systems EvaluationElisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren et al.SIGIR 2020 · 16 citations
