On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-n Recommendation
Olivier Jeunen, Ivan Potapov, Aleksei Ustimenko
摘要
Approaches to recommendation are typically evaluated in one of two ways: (1) via a (simulated) online experiment, often seen as the gold standard, or (2) via some offline evaluation procedure, where the goal is to approximate the outcome of an online experiment. Several offline evaluation metrics have been adopted in the literature, inspired by ranking metrics prevalent in the field of Information Retrieval. (Normalised) Discounted Cumulative Gain (nDCG) is one such metric that has seen widespread adoption in empirical studies, and higher (n)DCG values have been used to present new methods as the state-of-the-art in top-𝑛 recommendation for many years. Our work takes a critical look at this approach, and investigates when we can expect such metrics to approximate the gold standard outcome of an online experiment. We formally present the assumptions that are necessary to consider DCG an unbiased estimator of online reward and provide a derivation for this metric from first principles, highlighting where we deviate from its traditional uses in IR. Importantly, we show that normalising the metric renders it inconsistent, in that even when DCG is unbiased, ranking competing methods by their normalised DCG can invert their relative order. Through a correlation analysis between off-and on-line experiments conducted on a large-scale recommendation platform, we show that our unbiased DCG estimates strongly correlate with online reward, even when some of the metric's inherent assumptions are violated. This statement no longer holds for its normalised variant, suggesting that nDCG's practical utility may be limited. CCS CONCEPTS • Information systems → Recommender systems; Evaluation of retrieval results; • Mathematics of computing → Probabilistic inference problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning to Rank with Variable Result Presentation LengthsNorman Knyazev, Harrie OosterhuisSIGIR 2025 · 被引用 1 次
- Personalized Representation from Personalized GenerationShobhita Sundaram, Julia Chae, Yonglong Tian, Sara Beery 等ICLR 2025
它引用的顶会 Paper16
- On Sampled Metrics for Item RecommendationWalid Krichene, Steffen RendleKDD 2020 · 被引用 459 次
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 被引用 60 次
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 被引用 55 次
相关 Paper
- A Guided Learning Approach for Item Recommendation via Surrogate Loss LearningAhmed Rashed, Josif Grabocka, Lars Schmidt-ThiemeSIGIR 2021 · 被引用 10 次
- Large-scale Stochastic Optimization of NDCG Surrogates for Deep Learning with Provable ConvergenceZi-Hao Qiu, Quanqi Hu, Yongjian Zhong, Lijun Zhang 等ICML 2022 · 被引用 25 次
- New Insights into Metric Optimization for Ranking-based RecommendationRoger Zhe Li, Julián Urbano, Alan HanjalicSIGIR 2021 · 被引用 6 次
- Offline Retrieval Evaluation Without Evaluation MetricsFernando Diaz, Andres FerraroSIGIR 2022 · 被引用 8 次
- Agreement and Disagreement between True and False-Positive Metrics in Recommender Systems EvaluationElisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren 等SIGIR 2020 · 被引用 16 次
