A General Offline Reinforcement Learning Framework for Interactive Recommendation
Teng Xiao, Donglin Wang
摘要
This paper studies the problem of learning interactive recommender systems from logged feedbacks without any exploration in online environments. We address the problem by proposing a general offline reinforcement learning framework for recommendation, which enables maximizing cumulative user rewards without online exploration. Specifically, we first introduce a probabilistic generative model for interactive recommendation, and then propose an effective inference algorithm for discrete and stochastic policy learning based on logged feedbacks. In order to perform offline learning more effectively, we propose five approaches to minimize the distribution mismatch between the logging policy and recommendation policy: support constraints, supervised regularization, policy constraints, dual constraints and reward extrapolation. We conduct extensive experiments on two public real-world datasets, demonstrating that the proposed methods can achieve superior performance over existing supervised learning and reinforcement learning methods for recommendation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Elastic Decision TransformerYueh-Hua Wu, Xiaolong Wang, Masashi HamayaNeurIPS 2023 · 被引用 96 次
- Simple and Asymmetric Graph Contrastive Learning without AugmentationsTeng Xiao, Huaisheng Zhu, Zhengyu Chen, Suhang WangNeurIPS 2023 · 被引用 86 次
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li 等NeurIPS 2022 · 被引用 81 次
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li 等NeurIPS 2024 · 被引用 76 次
- Decoupled Self-supervised Learning for GraphsTeng Xiao, Zhengyu Chen, Zhimeng Guo, Zeyang Zhuang 等NeurIPS 2022 · 被引用 75 次
它引用的顶会 Paper5
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki 等ICLR 2020 · 被引用 299 次
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 被引用 217 次
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement LearningSijin Zhou, Xinyi Dai, Haokun Chen, Weinan Zhang 等SIGIR 2020 · 被引用 166 次
- Deep Transfer Tensor Decomposition with Orthogonal Constraint for Recommender SystemsZhengyu Chen, Ziqing Xu, Donglin WangAAAI 2021 · 被引用 53 次
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 被引用 22 次
相关 Paper
- Text-Based Interactive Recommendation via Offline Reinforcement LearningRuiyi Zhang, Tong Yu, Yilin Shen, Hongxia JinAAAI 2022 · 被引用 10 次
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li 等WWW 2025 · 被引用 4 次
- Advantage-Conditioned Flow Policy for Offline Reinforcement Learning in RecommendationXiaocong Chen, Siyu Wang, Lina YaoSIGIR 2026
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren 等SIGIR 2022 · 被引用 48 次
- Towards Off-Policy Learning for Ranking Policies with Logged FeedbackTeng Xiao, Suhang WangAAAI 2022 · 被引用 8 次
