Text-Based Interactive Recommendation via Offline Reinforcement Learning
Ruiyi Zhang, Tong Yu, Yilin Shen, Hongxia Jin
Abstract
Interactive recommendation with natural-language feedback can provide richer user feedback and has demonstrated advantages over traditional recommender systems. However, the classical online paradigm involves iteratively collecting experience via interaction with users, which is expensive and risky. We consider an offline interactive recommendation to exploit arbitrary experience collected by multiple unknown policies. A direct application of policy learning with such fixed experience suffers from the distribution shift. To tackle this issue, we develop a behavior-agnostic off-policy correction framework to make offline interactive recommendation possible. Specifically, we leverage the conservative Q-function to perform off-policy evaluation, which enables learning effective policies from fixed datasets without further interactions. Empirical results on the simulator derived from real-world datasets demonstrate the effectiveness of our proposed offline training framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Contrastive State Augmentations for Reinforcement Learning-Based Recommender SystemsZhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren et al.SIGIR 2023 · 20 citations
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng et al.WWW 2024 · 11 citations
- FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded MemoryAnwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang et al.ICCV 2023 · 11 citations
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
Builds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Interactive Path Reasoning on Graph for Conversational RecommendationWenqiang Lei, Gangyi Zhang, Xiangnan He, Yisong Miao et al.KDD 2020 · 158 citations
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang et al.WWW 2020 · 106 citations
Related papers
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
- Confidence-Conditioned Value Functions for Offline Reinforcement LearningJoey Hong, Aviral Kumar, Sergey LevineICLR 2023 · 4 citations
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 173 citations
- Behavior Prior Representation learning for Offline Reinforcement LearningHongyu Zang, Xin Li, Jie Yu, Chen Liu et al.ICLR 2023 · 3 citations
