Models Versus Satisfaction: Towards a Better Understanding of Evaluation Metrics
Fan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie, Weizhi Ma, Min Zhang, Shaoping Ma
Abstract
Evaluation metrics play an important role in the batch evaluation of IR systems. Based on a user model that describes how users interact with the rank list, an evaluation metric is defined to link the relevance scores of a list of documents to an estimation of system effectiveness and user satisfaction. Therefore, the validity of an evaluation metric has two facets: whether the underlying user model can accurately predict user behavior and whether the evaluation metric correlates well with user satisfaction. While a tremendous amount of work has been undertaken to design, evaluate, and compare different evaluation metrics, few studies have explored the consistency between these two facets of evaluation metrics. Specifically, we want to investigate whether the metrics that are well calibrated with user behavior data can perform as well in estimating user satisfaction. To shed light on this research question, we compare the performance of various metrics with the C/W/L Framework in estimating user satisfaction when they are optimized to fit observed user behavior. Experimental results on both self-collected and public available user search behavior datasets show that the metrics optimized to fit users' click behavior can perform as well as those calibrated with user satisfaction feedback. We also investigate the reliability in the calibration process of evaluation metrics to find out how much data is required for parameter tuning. Our findings provide empirical support for the consistency between user behavior modeling and satisfaction measurement, as well as guidance for tuning the parameters in evaluation metrics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 499bb0f1-224f-44d3-ac8e-fd9f48db6bb9Cited by top-tier papers4
- Towards a Better Understanding of Query Reformulation Behavior in Web SearchJia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang et al.WWW 2021 · 67 citations
- A Flexible Framework for Offline Effectiveness MetricsAlistair Moffat, Joel Mackenzie, Paul Thomas, Leif AzzopardiSIGIR 2022 · 41 citations
- Re-thinking Knowledge Graph Completion Evaluation from an Information Retrieval PerspectiveYing Zhou, Xuanang Chen, Ben He, Zheng Ye et al.SIGIR 2022 · 13 citations
- CROWN: A Novel Approach to Comprehending Users' Preferences for Accurate Personalized News RecommendationYunyong Ko, Seongeun Ryu, Sang-Wook KimWWW 2025 · 4 citations
Related papers
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 21 citations
- Why Don't You Click: Understanding Non-Click Results in Web Search with Brain SignalsZiyi Ye, Xiaohui Xie, Yiqun Liu, Zhihong Wang et al.SIGIR 2022 · 15 citations
- Global or Local: Constructing Personalized Click Models for Web SearchJunqi Zhang, Yiqun Liu, Jiaxin Mao, Xiaohui Xie et al.WWW 2022 · 6 citations
- Towards Better Evaluating Multi-Query Sessions: A Measure Based on the Theory of Planned BehaviorWenbo Zhang, Fan Zhang, Jia Chen, Wei LuSIGIR 2025 · 1 citation
- Preference-based Evaluation Metrics for Web Image SearchXiaohui Xie, Jiaxin Mao, Yiqun Liu, Maarten de Rijke et al.SIGIR 2020 · 12 citations
