Off-policy Learning over Heterogeneous Information for Recommendation
Xiangmeng Wang, Qian Li, Dianer Yu, Guandong Xu
摘要
Reinforcement learning has recently become an active topic in recommender system research, where the logged data that records interactions between items and users feedback is used to discover the policy. Much off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has been a popular research topic in reinforcement learning. However, the log entries are biased in that the logs over-represent actions favored by the recommender system, as the user feedback contains only partial information limited to the particular items exposed to the user. As a result, the policy learned from such off-line logged data tends to be biased from the true behaviour policy. In this paper, we are the first to propose a novel off-policy learning augmented by meta-paths for the recommendation. We argue that the Heterogeneous information network (HIN), which provides rich contextual information of items and user aspects, could scale the logged data contribution for unbiased target policy learning. Towards this end, we develop a new HIN augmented target policy model (HINpolicy), which explicitly leverages contextual information to scale the generated reward for target policy. In addition, being equipped with the HINpolicy model, our solution adaptively receives HIN-augmented corrections for counterfactual risk minimization, and ultimately yields an effective policy to maximize the long run rewards for the recommendation. Finally, we extensively evaluate our method through a series of simulations and large-scale real-world datasets, obtaining favorable results compared with state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Inductive Subgraphs as Shortcuts: Causal Disentanglement for Heterophilic Graph LearningXiangmeng Wang, Qian Li, Haiyang Xia, Hao Miao 等SIGIR 2026
- COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving GraphsSheng'en Li, Dongmian ZouICML 2026
它引用的顶会 Paper9
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 被引用 217 次
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
- Hierarchical Reinforcement Learning for Integrated RecommendationRuobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia 等AAAI 2021 · 被引用 90 次
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 被引用 82 次
相关 Paper
- MGPolicy: Meta Graph Enhanced Off-policy Learning for RecommendationsXiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang 等SIGIR 2022 · 被引用 8 次
- Meta-learning on Heterogeneous Information Networks for Cold-start RecommendationYuanfu Lu, Yuan Fang, Chuan ShiKDD 2020 · 被引用 255 次
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie 等NeurIPS 2023 · 被引用 6 次
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 被引用 24 次
- Reinforcement Learning Based Meta-Path Discovery in Large-Scale Heterogeneous Information NetworksGuojia Wan, Bo Du, Shirui Pan, Gholamreza HaffariAAAI 2020 · 被引用 45 次
