Evaluation Measures Based on Preference Graphs
Charles L. A. Clarke, Chengxi Luo, Mark D. Smucker
Abstract
The offline evaluation of search requires us to define a standard against which we measure the quality of results returned by a ranker. Frequently this standard is defined in absolute terms through relevance grades, but it can also be defined in relative terms through preferences. These preferences might be created through explicit preference judgments, derived from relevance grades, or inferred from clicks and other signals. Preferences from multiple sources might even be combined. In contrast to absolute grades, preferences avoid complex definitions of relevance, indicating only that a ranker should favor one result over another. Despite the simplicity and flexibility of preferences, widespread adoption has been limited by the lack of established evaluation measures. Recent work in this direction has taken two approaches: 1) measures based on weighted counts of agreements and disagreements between a set of preferences and an actual ranking generated by a ranker; and 2) measures that translate preferences into gain values for use with traditional measures, such as nDCG. Both approaches require methods for specifying weights or gains that have little or no theoretical foundation, and the values of these measures have no clear and meaningful interpretation. To address these problems, we propose an evaluation measure that computes the similarity between a directed multigraph of preferences and an actual ranking generated by a ranker. The measure computes an ordering for the vertices of the preference graph that maximizes its similarity to the actual ranking under a rank similarity measure. This maximum similarity becomes the value of the measure. Preference graphs are often acyclic, or nearly so, and to compute the measure we extend an approximate greedy algorithm that is known to produce good results for nearly acyclic graphs. For the rank similarity measure we employ Rank Biased Overlap (RBO) which was explicitly created to match the requirements of search and related applications. We validate the new measure over several collections of preferences explored in recent work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ba7dc54f-d69c-4768-b6ef-3b9cee987e6bCited by top-tier papers2
- Human Preferences as Dueling BanditsXinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell et al.SIGIR 2022 · 9 citations
- Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatEnrique Amigó, Stefano Mizzaro, Damiano SpinaSIGIR 2022 · 6 citations
Related papers
- The Treatment of Ties in Rank-Biased OverlapMatteo Corsi, Julián UrbanoSIGIR 2024 · 11 citations
- Good Evaluation Measures based on Document PreferencesTetsuya Sakai, Zhaohao ZengSIGIR 2020 · 14 citations
- Offline Retrieval Evaluation Without Evaluation MetricsFernando Diaz, Andres FerraroSIGIR 2022 · 8 citations
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 21 citations
- On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-n RecommendationOlivier Jeunen, Ivan Potapov, Aleksei UstimenkoKDD 2024 · 16 citations
