Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement Learning
Yang Deng, Yaliang Li, Fei Sun, Bolin Ding, Wai Lam
Abstract
Conversational recommender systems (CRS) enable the traditional recommender systems to explicitly acquire user preferences towards items and attributes through interactive conversations. Reinforcement learning (RL) is widely adopted to learn conversational recommendation policies to decide what attributes to ask, which items to recommend, and when to ask or recommend, at each conversation turn. However, existing methods mainly target at solving one or two of these three decision-making problems in CRS with separated conversation and recommendation components, which restrict the scalability and generality of CRS and fall short of preserving a stable training procedure. In the light of these challenges, we propose to formulate these three decision-making problems in CRS as a unified policy learning task. In order to systematically integrate conversation and recommendation components, we develop a dynamic weighted graph based RL method to learn a policy to select the action at each conversation turn, either asking an attribute or recommending items. Further, to deal with the sample efficiency issue, we propose two action selection strategies for reducing the candidate action space according to the preference and entropy information. Experimental results on two benchmark CRS datasets and a real-world E-Commerce application show that the proposed method not only significantly outperforms state-of-the-art methods but also enhances the scalability and stability of CRS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68594eed-32cf-430e-b97f-5e0c5d2b8592Cited by top-tier papers29
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 · 86 citations
- Multiple Choice Questions based Multi-Interest Policy Learning for Conversational RecommendationYiming Zhang, Lingfei Wu, Qi Shen, Yitong Pang et al.WWW 2022 · 72 citations
- Learning Neural Templates for Recommender Dialogue SystemZujie Liang, Huang Hu, Can Xu, Jian Miao et al.EMNLP 2021 · 40 citations
- User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational SystemsYang Deng, Wenxuan Zhang, Wai Lam, Hong Cheng et al.WWW 2022 · 34 citations
- Variational Reasoning about User Preferences for Conversational RecommendationZhaochun Ren, Zhi Tian, Dongdong Li, Pengjie Ren et al.SIGIR 2022 · 31 citations
Builds on11
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Improving Conversational Recommender Systems via Knowledge Graph based Semantic FusionKun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou et al.KDD 2020 · 309 citations
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Reinforced Negative Sampling over Knowledge Graph for RecommendationXiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao et al.WWW 2020 · 209 citations
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement LearningSijin Zhou, Xinyi Dai, Haokun Chen, Weinan Zhang et al.SIGIR 2020 · 166 citations
Related papers
- Confident Action Decision via Hierarchical Policy Learning for Conversational RecommendationHeeseon Kim, Hyeongjun Yang, Kyong-Ho LeeWWW 2023 · 10 citations
- Interactive Path Reasoning on Graph for Conversational RecommendationWenqiang Lei, Gangyi Zhang, Xiangnan He, Yisong Miao et al.KDD 2020 · 158 citations
- Learning to Infer User Implicit Preference in Conversational RecommendationChenhao Hu, Shuhua Huang, Yansen Zhang, Yubao LiuSIGIR 2022 · 38 citations
- Multi-Objective Intrinsic Reward Learning for Conversational Recommender SystemsZhendong Chu, Nan Wang, Hongning WangNeurIPS 2023 · 5 citations
- HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender SystemMingjie Qian, Yongsen Zheng, Jinghui Qin, Liang LinEMNLP 2023 · 11 citations
