Conversational Contextual Bandit: Algorithm and Application
Xiaoying Zhang, Hong Xie, Hang Li, John C. S. Lui
Abstract
Contextual bandit algorithms provide principled online learning solutions to balance the exploitation-exploration trade-off in various applications such as recommender systems. However, the learning speed of the traditional contextual bandit algorithms is often slow due to the need for extensive exploration. This poses a critical issue in applications like recommender systems, since users may need to provide feedbacks on a lot of uninterested items. To accelerate the learning speed, we generalize contextual bandit to conversational contextual bandit. Conversational contextual bandit leverages not only behavioral feedbacks on arms (e.g., articles in news recommendation), but also occasional conversational feedbacks on key-terms from the user. Here, a key-term can relate to a subset of arms, for example, a category of articles in news recommendation. We then design the Conversational UCB algorithm (ConUCB) to address two challenges in conversational contextual bandit: (1) which key-terms to select to conduct conversation, (2) how to leverage conversational feedbacks to accelerate the speed of bandit learning. We theoretically prove that ConUCB can achieve a smaller regret upper bound than the traditional contextual bandit algorithm LinUCB, which implies a faster learning speed. Experiments on synthetic data, as well as real datasets from Yelp and Toutiao, demonstrate the efficacy of the ConUCB algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7621008-b15c-44a6-9f87-a15bb273b3c8Cited by top-tier papers19
- Interactive Path Reasoning on Graph for Conversational RecommendationWenqiang Lei, Gangyi Zhang, Xiangnan He, Yisong Miao et al.KDD 2020 · 158 citations
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding et al.SIGIR 2021 · 131 citations
- CR-Walker: Tree-Structured Graph Reasoning and Dialog Acts for Conversational RecommendationWenchang Ma, Ryuichi Takanobu, Minlie HuangEMNLP 2021 · 45 citations
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 40 citations
- Variational Reasoning about User Preferences for Conversational RecommendationZhaochun Ren, Zhi Tian, Dongdong Li, Pengjie Ren et al.SIGIR 2022 · 31 citations
Related papers
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 18 citations
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual BanditsMaoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. LuiKDD 2025 · 1 citation
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang et al.KDD 2024 · 6 citations
- Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit LearningZhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. LuiWWW 2025 · 10 citations
- Syndicated Bandits: A Framework for Auto Tuning Hyper-parameters in Contextual Bandit AlgorithmsQin Ding, Yue Kang, Yi-Wei Liu, Thomas Chun Man Lee et al.NeurIPS 2022 · 12 citations
