MGPolicy: Meta Graph Enhanced Off-policy Learning for Recommendations
Xiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang, Hongxu Chen, Guandong Xu
Abstract
Off-policy learning has drawn huge attention in recommender systems (RS), which provides an opportunity for reinforcement learning to abandon the expensive online training. However, off-policy learning from logged data suffers biases caused by the policy shift between the target policy and the logging policy. Consequently, most off-policy learning resorts to inverse propensity scoring (IPS) which however tends to be over-fitted over exposed (or recommended) items and thus fails to explore unexposed items.
In this paper, we propose meta graph enhanced off-policy learning (MGPolicy), which is the first recommendation model for correcting the off-policy bias via contextual information. In particular, we explicitly leverage rich semantics in meta graphs for user state representation, and then train the candidate generation model to promote an efficient search in the action space. Moreover, our MGpolicy is designed with counterfactual risk minimization, which can correct policy learning bias and ultimately yield an effective target policy to maximize the long-run rewards for the recommendation. We extensively evaluate our method through a series of simulations and large-scale real-world datasets, achieving favorable results compared with state-of-the-art methods. Our code is currently available at https://www.dropbox.com/sh/9ugr1lx7gzwfub4/ AABY46hVG6qKJnGAWjRJZMFKa?dl=0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f9c9fbb-ed6d-4554-9c9d-da3fd84b1447Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement LearningSijin Zhou, Xinyi Dai, Haokun Chen, Weinan Zhang et al.SIGIR 2020 · 166 citations
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang et al.WWW 2020 · 106 citations
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
- Learning Causal Effects via Weighted Empirical Risk MinimizationYonghan Jung, Jin Tian, Elias BareinboimNeurIPS 2020 · 53 citations
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova et al.AAAI 2020 · 48 citations
Related papers
- Off-policy Learning over Heterogeneous Information for RecommendationXiangmeng Wang, Qian Li, Dianer Yu, Guandong XuWWW 2022 · 11 citations
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie et al.NeurIPS 2023 · 6 citations
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 24 citations
- Meta Graph Learning for Long-tail RecommendationChunyu Wei, Jian Liang, Di Liu, Zehui Dai et al.KDD 2023 · 22 citations
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 162 citations
