On Estimating Recommendation Evaluation Metrics under Sampling
Ruoming Jin, Dong Li, Benjamin Mudrak, Jing Gao, Zhi Liu
Abstract
Since the recent study (Krichene and Rendle 2020) done by Krichene and Rendle on the sampling-based top-k evaluation metric for recommendation, there has been a lot of debates on the validity of using sampling to evaluate recommendation algorithms. Though their work and the recent work (Li et al. 2020 ) have proposed some basic approaches for mapping the sampling-based metrics to their global counterparts which rank the entire set of items, there is still a lack of understanding and consensus on how sampling should be used for recommendation evaluation. The proposed approaches either are rather uninformative (linking sampling to metric evaluation) or can only work on simple metrics, such as Recall/Precision (Krichene and Rendle 2020; Li et al. 2020) . In this paper, we introduce a new research problem on learning the empirical rank distribution, and a new approach based on the estimated rank distribution, to estimate the top-k metrics. Since this question is closely related to the underlying mechanism of sampling for recommendation, tackling it can help better understand the power of sampling and can help resolve the questions of if and how should we use sampling for evaluating recommendation. We introduce two approaches based on MLE (Maximal Likelihood Estimation) and its weighted variants, and ME (Maximal Entropy) principals to recover the empirical rank distribution, and then utilize them for metrics estimation. The experimental results show the advantages of using the new approaches for evaluating recommendation algorithms based on top-k metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8461529d-a9f9-4eed-87eb-4e298fda21a6Cited by top-tier papers2
- Towards Reliable Item Sampling for Recommendation EvaluationDong Li, Ruoming Jin, Zhenming Liu, Bin Ren et al.AAAI 2023 · 11 citations
- IDTraffickers: An Authorship Attribution Dataset to link and connect Potential Human-Trafficking Operations on Text Escort AdvertisementsVageesh Saxena, Benjamin Bashpole, Gijs van Dijck, Gerasimos SpanakisEMNLP 2023 · 2 citations
Builds on3
- On Sampled Metrics for Item RecommendationWalid Krichene, Steffen RendleKDD 2020 · 459 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
- On Sampling Top-K Recommendation EvaluationDong Li, Ruoming Jin, Jing Gao, Zhi LiuKDD 2020 · 43 citations
Related papers
- New Insights into Metric Optimization for Ranking-based RecommendationRoger Zhe Li, Julián Urbano, Alan HanjalicSIGIR 2021 · 6 citations
- Talos: Optimizing Top-K Accuracy in Recommender SystemsShengjia Zhang, Weiqin Yang, Jiawei Chen, Peng Wu et al.WWW 2026 · 1 citation
- On the Theories Behind Hard Negative Sampling for RecommendationWentao Shi, Jiawei Chen, Fuli Feng, Jizhi Zhang et al.WWW 2023 · 66 citations
- Weighted Sampling without Replacement for Deep Top-k ClassificationDieqiao Feng, Yuanqi Du, Carla P. Gomes, Bart SelmanICML 2023
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 60 citations
