Off-policy Learning over Heterogeneous Information for Recommendation
Xiangmeng Wang, Qian Li, Dianer Yu, Guandong Xu
Abstract
Reinforcement learning has recently become an active topic in recommender system research, where the logged data that records interactions between items and users feedback is used to discover the policy. Much off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has been a popular research topic in reinforcement learning. However, the log entries are biased in that the logs over-represent actions favored by the recommender system, as the user feedback contains only partial information limited to the particular items exposed to the user. As a result, the policy learned from such off-line logged data tends to be biased from the true behaviour policy. In this paper, we are the first to propose a novel off-policy learning augmented by meta-paths for the recommendation. We argue that the Heterogeneous information network (HIN), which provides rich contextual information of items and user aspects, could scale the logged data contribution for unbiased target policy learning. Towards this end, we develop a new HIN augmented target policy model (HINpolicy), which explicitly leverages contextual information to scale the generated reward for target policy. In addition, being equipped with the HINpolicy model, our solution adaptively receives HIN-augmented corrections for counterfactual risk minimization, and ultimately yields an effective policy to maximize the long run rewards for the recommendation. Finally, we extensively evaluate our method through a series of simulations and large-scale real-world datasets, obtaining favorable results compared with state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac02ca0b-be08-4861-a004-c2d5f8ff2a2aCited by top-tier papers2
- Inductive Subgraphs as Shortcuts: Causal Disentanglement for Heterophilic Graph LearningXiangmeng Wang, Qian Li, Haiyang Xia, Hao Miao et al.SIGIR 2026
- COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving GraphsSheng'en Li, Dongmian ZouICML 2026
Builds on9
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang et al.WWW 2020 · 106 citations
- Hierarchical Reinforcement Learning for Integrated RecommendationRuobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia et al.AAAI 2021 · 90 citations
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
Related papers
- MGPolicy: Meta Graph Enhanced Off-policy Learning for RecommendationsXiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang et al.SIGIR 2022 · 8 citations
- Meta-learning on Heterogeneous Information Networks for Cold-start RecommendationYuanfu Lu, Yuan Fang, Chuan ShiKDD 2020 · 255 citations
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie et al.NeurIPS 2023 · 6 citations
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 24 citations
- Reinforcement Learning Based Meta-Path Discovery in Large-Scale Heterogeneous Information NetworksGuojia Wan, Bo Du, Shirui Pan, Gholamreza HaffariAAAI 2020 · 45 citations
