MGPolicy: Meta Graph Enhanced Off-policy Learning for Recommendations
Xiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang, Hongxu Chen, Guandong Xu
摘要
Off-policy learning has drawn huge attention in recommender systems (RS), which provides an opportunity for reinforcement learning to abandon the expensive online training. However, off-policy learning from logged data suffers biases caused by the policy shift between the target policy and the logging policy. Consequently, most off-policy learning resorts to inverse propensity scoring (IPS) which however tends to be over-fitted over exposed (or recommended) items and thus fails to explore unexposed items.
In this paper, we propose meta graph enhanced off-policy learning (MGPolicy), which is the first recommendation model for correcting the off-policy bias via contextual information. In particular, we explicitly leverage rich semantics in meta graphs for user state representation, and then train the candidate generation model to promote an efficient search in the action space. Moreover, our MGpolicy is designed with counterfactual risk minimization, which can correct policy learning bias and ultimately yield an effective target policy to maximize the long-run rewards for the recommendation. We extensively evaluate our method through a series of simulations and large-scale real-world datasets, achieving favorable results compared with state-of-the-art methods. Our code is currently available at https://www.dropbox.com/sh/9ugr1lx7gzwfub4/ AABY46hVG6qKJnGAWjRJZMFKa?dl=0
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement LearningSijin Zhou, Xinyi Dai, Haokun Chen, Weinan Zhang 等SIGIR 2020 · 被引用 166 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 被引用 82 次
- Learning Causal Effects via Weighted Empirical Risk MinimizationYonghan Jung, Jin Tian, Elias BareinboimNeurIPS 2020 · 被引用 53 次
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova 等AAAI 2020 · 被引用 48 次
相关 Paper
- Off-policy Learning over Heterogeneous Information for RecommendationXiangmeng Wang, Qian Li, Dianer Yu, Guandong XuWWW 2022 · 被引用 11 次
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie 等NeurIPS 2023 · 被引用 6 次
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 被引用 24 次
- Meta Graph Learning for Long-tail RecommendationChunyu Wei, Jian Liang, Di Liu, Zehui Dai 等KDD 2023 · 被引用 22 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
