A General Offline Reinforcement Learning Framework for Interactive Recommendation
Teng Xiao, Donglin Wang
Abstract
This paper studies the problem of learning interactive recommender systems from logged feedbacks without any exploration in online environments. We address the problem by proposing a general offline reinforcement learning framework for recommendation, which enables maximizing cumulative user rewards without online exploration. Specifically, we first introduce a probabilistic generative model for interactive recommendation, and then propose an effective inference algorithm for discrete and stochastic policy learning based on logged feedbacks. In order to perform offline learning more effectively, we propose five approaches to minimize the distribution mismatch between the logging policy and recommendation policy: support constraints, supervised regularization, policy constraints, dual constraints and reward extrapolation. We conduct extensive experiments on two public real-world datasets, demonstrating that the proposed methods can achieve superior performance over existing supervised learning and reinforcement learning methods for recommendation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e14918e-d411-40a7-a659-af5d54cd79edCited by top-tier papers23
- Elastic Decision TransformerYueh-Hua Wu, Xiaolong Wang, Masashi HamayaNeurIPS 2023 · 96 citations
- Simple and Asymmetric Graph Contrastive Learning without AugmentationsTeng Xiao, Huaisheng Zhu, Zhengyu Chen, Suhang WangNeurIPS 2023 · 86 citations
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li et al.NeurIPS 2024 · 76 citations
- Decoupled Self-supervised Learning for GraphsTeng Xiao, Zhengyu Chen, Zhimeng Guo, Zeyang Zhuang et al.NeurIPS 2022 · 75 citations
Builds on5
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki et al.ICLR 2020 · 299 citations
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement LearningSijin Zhou, Xinyi Dai, Haokun Chen, Weinan Zhang et al.SIGIR 2020 · 166 citations
- Deep Transfer Tensor Decomposition with Orthogonal Constraint for Recommender SystemsZhengyu Chen, Ziqing Xu, Donglin WangAAAI 2021 · 53 citations
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 22 citations
Related papers
- Text-Based Interactive Recommendation via Offline Reinforcement LearningRuiyi Zhang, Tong Yu, Yilin Shen, Hongxia JinAAAI 2022 · 10 citations
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li et al.WWW 2025 · 4 citations
- Advantage-Conditioned Flow Policy for Offline Reinforcement Learning in RecommendationXiaocong Chen, Siyu Wang, Lina YaoSIGIR 2026
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
- Towards Off-Policy Learning for Ranking Policies with Logged FeedbackTeng Xiao, Suhang WangAAAI 2022 · 8 citations
