Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit Learning
Zhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. Lui
摘要
In conversational recommender systems, interactively presenting queries and leveraging user feedback are crucial for efficiently estimating user preferences and improving recommendation quality. Selecting optimal queries in these systems is a significant challenge that has been extensively studied as a sequential decision problem. The expected value of information (EVOI), which computes the expected reward improvement, provides a principled criterion for query selection. However, it is computationally expensive and lacks theoretical performance guarantees. Conversely, conversational bandits offer provable regret upper bounds, but their query selection strategies yield only marginal regret improvements over non-conversational approaches. To address these limitations, we integrate EVOI within the conversational bandit framework by proposing a new conversational mechanism featuring two key techniques: (1) gradient-based EVOI, which replaces the complex Bayesian updates in conventional EVOI with efficient stochastic gradient descent, significantly reducing computational complexity and facilitating theoretical analysis; and (2) smoothed key term contexts, which enhance exploration by adding random perturbations to uncover more specific user preferences. Our approach applies to both Bayesian (Thompson Sampling) and frequentist (UCB) variants of conversational bandits. We introduce two new algorithms, ConTS-EVOI and ConUCB-EVOI, and rigorously prove that they achieve substantially tighter regret bounds, with both algorithms offering a √ 𝑑 improvement in their dependence on the time horizon 𝑇 , where 𝑑 is the dimension of the feature space. Extensive evaluations on synthetic and real-world datasets validate the effectiveness of our methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Unified Online-Offline Framework for Co-Branding Campaign RecommendationsXiangxiang Dai, Xiaowei Sun, Jinhang Zuo, Xutong Liu 等KDD 2025 · 被引用 1 次
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual BanditsMaoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. LuiKDD 2025 · 被引用 1 次
- Not All Information Brings Benefits: Personalization-Driven Agent Debate for Conversational RecommendationPengfei Zhang, Guojia An, Jin Huang, Yuhan Yang 等WWW 2026
- Trading Vector Data in Vector DatabasesJin Cheng, Xiangxiang Dai, Ningning Ding, John C. S. Lui 等ICDE 2026
它引用的顶会 Paper8
- Conversational Contextual Bandit: Algorithm and ApplicationXiaoying Zhang, Hong Xie, Hang Li, John C. S. LuiWWW 2020 · 被引用 97 次
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 被引用 40 次
- Gradient-Based Optimization for Bayesian Preference ElicitationIvan Vendrov, Tyler Lu, Qingqing Huang, Craig BoutilierAAAI 2020 · 被引用 29 次
- Structured Linear Contextual Bandits: A Sharp and Geometric Smoothed AnalysisVidyashankar Sivakumar, Zhiwei Steven Wu, Arindam BanerjeeICML 2020 · 被引用 24 次
- Knowledge-aware Conversational Preference Elicitation with Bandit FeedbackCanzhe Zhao, Tong Yu, Zhihui Xie, Shuai LiWWW 2022 · 被引用 24 次
相关 Paper
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 被引用 18 次
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang 等KDD 2024 · 被引用 6 次
- Expected Improvement for Contextual BanditsHung Tran-The, Sunil Gupta, Santu Rana, Tuan Truong 等NeurIPS 2022 · 被引用 5 次
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding 等SIGIR 2021 · 被引用 131 次
- Learning to Infer User Implicit Preference in Conversational RecommendationChenhao Hu, Shuhua Huang, Yansen Zhang, Yubao LiuSIGIR 2022 · 被引用 38 次
