Conversational Contextual Bandit: Algorithm and Application
Xiaoying Zhang, Hong Xie, Hang Li, John C. S. Lui
摘要
Contextual bandit algorithms provide principled online learning solutions to balance the exploitation-exploration trade-off in various applications such as recommender systems. However, the learning speed of the traditional contextual bandit algorithms is often slow due to the need for extensive exploration. This poses a critical issue in applications like recommender systems, since users may need to provide feedbacks on a lot of uninterested items. To accelerate the learning speed, we generalize contextual bandit to conversational contextual bandit. Conversational contextual bandit leverages not only behavioral feedbacks on arms (e.g., articles in news recommendation), but also occasional conversational feedbacks on key-terms from the user. Here, a key-term can relate to a subset of arms, for example, a category of articles in news recommendation. We then design the Conversational UCB algorithm (ConUCB) to address two challenges in conversational contextual bandit: (1) which key-terms to select to conduct conversation, (2) how to leverage conversational feedbacks to accelerate the speed of bandit learning. We theoretically prove that ConUCB can achieve a smaller regret upper bound than the traditional contextual bandit algorithm LinUCB, which implies a faster learning speed. Experiments on synthetic data, as well as real datasets from Yelp and Toutiao, demonstrate the efficacy of the ConUCB algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Interactive Path Reasoning on Graph for Conversational RecommendationWenqiang Lei, Gangyi Zhang, Xiangnan He, Yisong Miao 等KDD 2020 · 被引用 158 次
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding 等SIGIR 2021 · 被引用 131 次
- CR-Walker: Tree-Structured Graph Reasoning and Dialog Acts for Conversational RecommendationWenchang Ma, Ryuichi Takanobu, Minlie HuangEMNLP 2021 · 被引用 45 次
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 被引用 40 次
- Variational Reasoning about User Preferences for Conversational RecommendationZhaochun Ren, Zhi Tian, Dongdong Li, Pengjie Ren 等SIGIR 2022 · 被引用 31 次
相关 Paper
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 被引用 18 次
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual BanditsMaoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. LuiKDD 2025 · 被引用 1 次
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang 等KDD 2024 · 被引用 6 次
- Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit LearningZhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. LuiWWW 2025 · 被引用 10 次
- Syndicated Bandits: A Framework for Auto Tuning Hyper-parameters in Contextual Bandit AlgorithmsQin Ding, Yue Kang, Yi-Wei Liu, Thomas Chun Man Lee 等NeurIPS 2022 · 被引用 12 次
