Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual Bandits
Maoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. Lui
摘要
Conversational recommender systems proactively query users with relevant "key terms" and leverage the feedback to elicit users' preferences for personalized recommendations. Conversational contextual bandits, a prevalent approach in this domain, aim to optimize preference learning by balancing exploitation and exploration. However, several limitations hinder their effectiveness in real-world scenarios. First, existing algorithms employ key term selection strategies with insufficient exploration, often failing to thoroughly probe users' preferences and resulting in suboptimal preference estimation. Second, current algorithms typically rely on deterministic rules to initiate conversations, causing unnecessary interactions when preferences are well-understood and missed opportunities when preferences are uncertain. To address these limitations, we propose three novel algorithms: CLiSK, CLiME, and CLiSK-ME. CLiSK introduces smoothed key term contexts to enhance exploration in preference learning, CLiME adaptively initiates conversations based on preference uncertainty, and CLiSK-ME integrates both techniques. We theoretically prove that all three algorithms achieve a tighter regret upper bound of O ( √︁ 𝑑𝑇 log𝑇 ) with respect to the time horizon 𝑇 , improving upon existing methods. Additionally, we provide a matching lower bound Ω( √ 𝑑𝑇 ) for conversational bandits, demonstrating that our algorithms are nearly minimax optimal. Extensive evaluations on both synthetic and real-world datasets show that our approaches achieve at least a 14.6% improvement in cumulative regret.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Conversational Contextual Bandit: Algorithm and ApplicationXiaoying Zhang, Hong Xie, Hang Li, John C. S. LuiWWW 2020 · 被引用 97 次
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 被引用 40 次
- Structured Linear Contextual Bandits: A Sharp and Geometric Smoothed AnalysisVidyashankar Sivakumar, Zhiwei Steven Wu, Arindam BanerjeeICML 2020 · 被引用 24 次
- Knowledge-aware Conversational Preference Elicitation with Bandit FeedbackCanzhe Zhao, Tong Yu, Zhihui Xie, Shuai LiWWW 2022 · 被引用 24 次
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 被引用 18 次
相关 Paper
- Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit LearningZhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. LuiWWW 2025 · 被引用 10 次
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang 等KDD 2024 · 被引用 6 次
- Hierarchical Adaptive Contextual Bandits for Resource Constraint based RecommendationMengyue Yang, Qingyang Li, Zhiwei (Tony) Qin, Jieping YeWWW 2020 · 被引用 13 次
- Online Clustering of Bandits with Misspecified User ModelsZhiyong Wang, Jize Xie, Xutong Liu, Shuai Li 等NeurIPS 2023 · 被引用 16 次
- Syndicated Bandits: A Framework for Auto Tuning Hyper-parameters in Contextual Bandit AlgorithmsQin Ding, Yue Kang, Yi-Wei Liu, Thomas Chun Man Lee 等NeurIPS 2022 · 被引用 12 次
