Models Versus Satisfaction: Towards a Better Understanding of Evaluation Metrics
Fan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie, Weizhi Ma, Min Zhang, Shaoping Ma
摘要
Evaluation metrics play an important role in the batch evaluation of IR systems. Based on a user model that describes how users interact with the rank list, an evaluation metric is defined to link the relevance scores of a list of documents to an estimation of system effectiveness and user satisfaction. Therefore, the validity of an evaluation metric has two facets: whether the underlying user model can accurately predict user behavior and whether the evaluation metric correlates well with user satisfaction. While a tremendous amount of work has been undertaken to design, evaluate, and compare different evaluation metrics, few studies have explored the consistency between these two facets of evaluation metrics. Specifically, we want to investigate whether the metrics that are well calibrated with user behavior data can perform as well in estimating user satisfaction. To shed light on this research question, we compare the performance of various metrics with the C/W/L Framework in estimating user satisfaction when they are optimized to fit observed user behavior. Experimental results on both self-collected and public available user search behavior datasets show that the metrics optimized to fit users' click behavior can perform as well as those calibrated with user satisfaction feedback. We also investigate the reliability in the calibration process of evaluation metrics to find out how much data is required for parameter tuning. Our findings provide empirical support for the consistency between user behavior modeling and satisfaction measurement, as well as guidance for tuning the parameters in evaluation metrics.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Towards a Better Understanding of Query Reformulation Behavior in Web SearchJia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang 等WWW 2021 · 被引用 67 次
- A Flexible Framework for Offline Effectiveness MetricsAlistair Moffat, Joel Mackenzie, Paul Thomas, Leif AzzopardiSIGIR 2022 · 被引用 41 次
- Re-thinking Knowledge Graph Completion Evaluation from an Information Retrieval PerspectiveYing Zhou, Xuanang Chen, Ben He, Zheng Ye 等SIGIR 2022 · 被引用 13 次
- CROWN: A Novel Approach to Comprehending Users' Preferences for Accurate Personalized News RecommendationYunyong Ko, Seongeun Ryu, Sang-Wook KimWWW 2025 · 被引用 4 次
相关 Paper
- A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational UsersNuo Chen, Jiqun Liu, Tetsuya SakaiWWW 2023 · 被引用 21 次
- Why Don't You Click: Understanding Non-Click Results in Web Search with Brain SignalsZiyi Ye, Xiaohui Xie, Yiqun Liu, Zhihong Wang 等SIGIR 2022 · 被引用 15 次
- Global or Local: Constructing Personalized Click Models for Web SearchJunqi Zhang, Yiqun Liu, Jiaxin Mao, Xiaohui Xie 等WWW 2022 · 被引用 6 次
- Towards Better Evaluating Multi-Query Sessions: A Measure Based on the Theory of Planned BehaviorWenbo Zhang, Fan Zhang, Jia Chen, Wei LuSIGIR 2025 · 被引用 1 次
- Preference-based Evaluation Metrics for Web Image SearchXiaohui Xie, Jiaxin Mao, Yiqun Liu, Maarten de Rijke 等SIGIR 2020 · 被引用 12 次
