Joint Policy-Value Learning for Recommendation
Olivier Jeunen, David Rohde, Flavian Vasile, Martin Bompaire
摘要
Conventional approaches to recommendation often do not explicitly take into account information on previously shown recommendations and their recorded responses. One reason is that, since we do not know the outcome of actions the system did not take, learning directly from such logs is not a straightforward task. Several methods for off-policy or counterfactual learning have been proposed in recent years, but their efficacy for the recommendation task remains understudied. Due to the limitations of offline datasets and the lack of access of most academic researchers to online experiments, this is a non-trivial task. Simulation environments can provide a reproducible solution to this problem. In this work, we conduct the first broad empirical study of counterfactual learning methods for recommendation, in a simulated environment. We consider various different policy-based methods that make use of the Inverse Propensity Score (IPS) to perform Counterfactual Risk Minimisation (CRM), as well as value-based methods based on Maximum Likelihood Estimation (MLE). We highlight how existing off-policy learning methods fail due to stochastic and sparse rewards, and show how a logarithmic variant of the traditional IPS estimator can solve these issues, whilst convexifying the objective and thus facilitating its optimisation. Additionally, under certain assumptions the value-and policy-based methods have an identical parameterisation, allowing us to propose a new model that combines both the MLE and CRM objectives. Extensive experiments show that this łDual Banditž approach achieves stateof-the-art performance in a wide range of scenarios, for varying logging policies, action spaces and training sample sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BLOB: A Probabilistic Model for Recommendation that Combines Organic and Bandit SignalsOtmane Sakhi, Stephen Bonner, David Rohde, Flavian VasileKDD 2020 · 被引用 18 次
- On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-n RecommendationOlivier Jeunen, Ivan Potapov, Aleksei UstimenkoKDD 2024 · 被引用 16 次
- Off-policy Learning over Heterogeneous Information for RecommendationXiangmeng Wang, Qian Li, Dianer Yu, Guandong XuWWW 2022 · 被引用 11 次
- AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution GeneralizationXiao Zhang, Sunhao Dai, Jun Xu, Yong Liu 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper2
相关 Paper
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie 等NeurIPS 2023 · 被引用 6 次
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 被引用 22 次
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging PoliciesKoichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita 等ICLR 2026 · 被引用 1 次
- MGPolicy: Meta Graph Enhanced Off-policy Learning for RecommendationsXiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang 等SIGIR 2022 · 被引用 8 次
- Off-Policy Evaluation of Ranking Policies under Diverse User BehaviorHaruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu 等KDD 2023 · 被引用 8 次
