A Flexible Framework for Offline Effectiveness Metrics
Alistair Moffat, Joel Mackenzie, Paul Thomas, Leif Azzopardi
摘要
The use of offline effectiveness metrics is one of the cornerstones of evaluation in information retrieval. Static resources that include test collections and sets of topics, the corresponding relevance judgments connecting them, and metrics that map document rankings from a retrieval system to numeric scores have been used for multiple decades as an important way of comparing systems. The basis behind this experimental structure is that the metric score for a system can serve as a surrogate measurement for user satisfaction.
Here we introduce a user behavior framework that extends the C/W/L family. The essence of the new framework -which we call C/W/L/A -is that the user actions that are undertaken while reading the ranking can be considered separately from the benefit that each user will have derived as they exit the ranking. This split structure allows the great majority of current effectiveness metrics to be systematically categorized, and thus their relative properties and relationships to be better understood; and at the same time permits a wide range of novel combinations to be considered.
We then carry out experiments using relevance judgments, document rankings, and user satisfaction data from two distinct sources, comparing the patterns of metric scores generated, and showing that those metrics vary quite markedly in terms of their ability to predict user satisfaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatEnrique Amigó, Stefano Mizzaro, Damiano SpinaSIGIR 2022 · 被引用 6 次
- Offline Evaluation of Ranked Lists using Parametric Estimation of PropensitiesVishwa Vinay, Manoj Kilaru, David ArbourSIGIR 2022
- LLM-as-a-Judge for Reliable and Explainable Offline Evaluation in Top-K RecommendationYue Que, Junyi Zhou, Xiaokun Zhang, Haiming Jin 等KDD 2026
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 被引用 21 次
- Why Don't You Click: Understanding Non-Click Results in Web Search with Brain SignalsZiyi Ye, Xiaohui Xie, Yiqun Liu, Zhihong Wang 等SIGIR 2022 · 被引用 15 次
