A Flexible Framework for Offline Effectiveness Metrics
Alistair Moffat, Joel Mackenzie, Paul Thomas, Leif Azzopardi
Abstract
The use of offline effectiveness metrics is one of the cornerstones of evaluation in information retrieval. Static resources that include test collections and sets of topics, the corresponding relevance judgments connecting them, and metrics that map document rankings from a retrieval system to numeric scores have been used for multiple decades as an important way of comparing systems. The basis behind this experimental structure is that the metric score for a system can serve as a surrogate measurement for user satisfaction.
Here we introduce a user behavior framework that extends the C/W/L family. The essence of the new framework -which we call C/W/L/A -is that the user actions that are undertaken while reading the ranking can be considered separately from the benefit that each user will have derived as they exit the ranking. This split structure allows the great majority of current effectiveness metrics to be systematically categorized, and thus their relative properties and relationships to be better understood; and at the same time permits a wide range of novel combinations to be considered.
We then carry out experiments using relevance judgments, document rankings, and user satisfaction data from two distinct sources, comparing the patterns of metric scores generated, and showing that those metrics vary quite markedly in terms of their ability to predict user satisfaction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d7cd9c3-2ad6-437f-b0cd-8b9d133738aaBuilds on2
- Towards a Better Understanding of Query Reformulation Behavior in Web SearchJia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang et al.WWW 2021 · 67 citations
- Models Versus Satisfaction: Towards a Better Understanding of Evaluation MetricsFan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie et al.SIGIR 2020 · 35 citations
Related papers
- Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatEnrique Amigó, Stefano Mizzaro, Damiano SpinaSIGIR 2022 · 6 citations
- Offline Evaluation of Ranked Lists using Parametric Estimation of PropensitiesVishwa Vinay, Manoj Kilaru, David ArbourSIGIR 2022
- LLM-as-a-Judge for Reliable and Explainable Offline Evaluation in Top-K RecommendationYue Que, Junyi Zhou, Xiaokun Zhang, Haiming Jin et al.KDD 2026
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 21 citations
- Why Don't You Click: Understanding Non-Click Results in Web Search with Brain SignalsZiyi Ye, Xiaohui Xie, Yiqun Liu, Zhihong Wang et al.SIGIR 2022 · 15 citations
