Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit Learning
Zhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. Lui
Abstract
In conversational recommender systems, interactively presenting queries and leveraging user feedback are crucial for efficiently estimating user preferences and improving recommendation quality. Selecting optimal queries in these systems is a significant challenge that has been extensively studied as a sequential decision problem. The expected value of information (EVOI), which computes the expected reward improvement, provides a principled criterion for query selection. However, it is computationally expensive and lacks theoretical performance guarantees. Conversely, conversational bandits offer provable regret upper bounds, but their query selection strategies yield only marginal regret improvements over non-conversational approaches. To address these limitations, we integrate EVOI within the conversational bandit framework by proposing a new conversational mechanism featuring two key techniques: (1) gradient-based EVOI, which replaces the complex Bayesian updates in conventional EVOI with efficient stochastic gradient descent, significantly reducing computational complexity and facilitating theoretical analysis; and (2) smoothed key term contexts, which enhance exploration by adding random perturbations to uncover more specific user preferences. Our approach applies to both Bayesian (Thompson Sampling) and frequentist (UCB) variants of conversational bandits. We introduce two new algorithms, ConTS-EVOI and ConUCB-EVOI, and rigorously prove that they achieve substantially tighter regret bounds, with both algorithms offering a √ 𝑑 improvement in their dependence on the time horizon 𝑇 , where 𝑑 is the dimension of the feature space. Extensive evaluations on synthetic and real-world datasets validate the effectiveness of our methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe26deb1-9282-44a0-a76a-3757e089bc1bCited by top-tier papers4
- A Unified Online-Offline Framework for Co-Branding Campaign RecommendationsXiangxiang Dai, Xiaowei Sun, Jinhang Zuo, Xutong Liu et al.KDD 2025 · 1 citation
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual BanditsMaoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. LuiKDD 2025 · 1 citation
- Not All Information Brings Benefits: Personalization-Driven Agent Debate for Conversational RecommendationPengfei Zhang, Guojia An, Jin Huang, Yuhan Yang et al.WWW 2026
- Trading Vector Data in Vector DatabasesJin Cheng, Xiangxiang Dai, Ningning Ding, John C. S. Lui et al.ICDE 2026
Builds on8
- Conversational Contextual Bandit: Algorithm and ApplicationXiaoying Zhang, Hong Xie, Hang Li, John C. S. LuiWWW 2020 · 97 citations
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 40 citations
- Gradient-Based Optimization for Bayesian Preference ElicitationIvan Vendrov, Tyler Lu, Qingqing Huang, Craig BoutilierAAAI 2020 · 29 citations
- Structured Linear Contextual Bandits: A Sharp and Geometric Smoothed AnalysisVidyashankar Sivakumar, Zhiwei Steven Wu, Arindam BanerjeeICML 2020 · 24 citations
- Knowledge-aware Conversational Preference Elicitation with Bandit FeedbackCanzhe Zhao, Tong Yu, Zhihui Xie, Shuai LiWWW 2022 · 24 citations
Related papers
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 18 citations
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang et al.KDD 2024 · 6 citations
- Expected Improvement for Contextual BanditsHung Tran-The, Sunil Gupta, Santu Rana, Tuan Truong et al.NeurIPS 2022 · 5 citations
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding et al.SIGIR 2021 · 131 citations
- Learning to Infer User Implicit Preference in Conversational RecommendationChenhao Hu, Shuhua Huang, Yansen Zhang, Yubao LiuSIGIR 2022 · 38 citations
