Offline Retrieval Evaluation Without Evaluation Metrics
Fernando Diaz, Andres Ferraro
摘要
Offline evaluation of information retrieval and recommendation has traditionally focused on distilling the quality of a ranking into a scalar metric such as average precision or normalized discounted cumulative gain. We can use this metric to compare the performance of multiple systems for the same request. Although evaluation metrics provide a convenient summary of system performance, they also collapse subtle differences across users into a single number and can carry assumptions about user behavior and utility not supported across retrieval scenarios. We propose recall-paired preference (RPP), a metric-free evaluation method based on directly computing a preference between ranked lists. RPP simulates multiple user subpopulations per query and compares systems across these pseudo-populations. Our results across multiple search and recommendation tasks demonstrate that RPP substantially improves discriminative power while correlating well with existing metrics and being equally robust to incomplete data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Evaluation Measures Based on Preference GraphsCharles L. A. Clarke, Chengxi Luo, Mark D. SmuckerSIGIR 2021 · 被引用 5 次
- Good Evaluation Measures based on Document PreferencesTetsuya Sakai, Zhaohao ZengSIGIR 2020 · 被引用 14 次
- Offline Evaluation of Ranked Lists using Parametric Estimation of PropensitiesVishwa Vinay, Manoj Kilaru, David ArbourSIGIR 2022
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 被引用 21 次
- Agreement and Disagreement between True and False-Positive Metrics in Recommender Systems EvaluationElisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren 等SIGIR 2020 · 被引用 16 次
