Conversational Dueling Bandits in Generalized Linear Models
Shuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang, Hong Zhang, Huazheng Wang
Abstract
Conversational recommendation systems elicit user preferences by interacting with users to obtain their feedback on recommended commodities. Such systems utilize a multi-armed bandit framework to learn user preferences in an online manner and have received great success in recent years. However, existing conversational bandit methods have several limitations. First, they only enable users to provide explicit binary feedback on the recommended items or categories, leading to ambiguity in interpretation. In practice, users are usually faced with more than one choice. Relative feedback, known for its informativeness, has gained increasing popularity in recommendation system design. Moreover, current contextual bandit methods mainly work under linear reward assumptions, ignoring practical non-linear reward structures in generalized linear models. Therefore, in this paper, we introduce relative feedbackbased conversations into conversational recommendation systems through the integration of dueling bandits in generalized linear models (GLM) and propose a novel conversational dueling bandit algorithm called ConDuel. Theoretical analyses of regret upper bounds and empirical validations on synthetic and real-world data underscore ConDuel's efficacy. We also demonstrate the potential to extend our algorithm to multinomial logit bandits with theoretical and experimental guarantees, which further proves the applicability of the proposed framework. CCS Concepts • Theory of computation → Online algorithms; • Information systems → Recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dee46ff2-94e0-49fc-bb11-688d99d1a3afCited by top-tier papers4
- Towards Efficient Conversational Recommendations: Expected Value of Information Meets Bandit LearningZhuohua Li, Maoli Liu, Xiangxiang Dai, John C. S. LuiWWW 2025 · 10 citations
- Large Language Model-Enhanced Multi-Armed BanditsJiahang Sun, Zhiyong Wang, Runhan Yang, Chenjun Xiao et al.ACL 2026 · 6 citations
- A Unified Online-Offline Framework for Co-Branding Campaign RecommendationsXiangxiang Dai, Xiaowei Sun, Jinhang Zuo, Xutong Liu et al.KDD 2025 · 1 citation
- Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual BanditsMaoli Liu, Zhuohua Li, Xiangxiang Dai, John C. S. LuiKDD 2025 · 1 citation
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Conversational Contextual Bandit: Algorithm and ApplicationXiaoying Zhang, Hong Xie, Hang Li, John C. S. LuiWWW 2020 · 97 citations
- Preference-based Reinforcement Learning with Finite-Time GuaranteesYichong Xu, Ruosong Wang, Lin F. Yang, Aarti Singh et al.NeurIPS 2020 · 82 citations
- Optimal Algorithms for Stochastic Contextual Preference BanditsAadirupa SahaNeurIPS 2021 · 64 citations
Related papers
- Online Clustering of Dueling BanditsZhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie et al.ICML 2025
- Comparison-based Conversational Recommender System with Relative Bandit FeedbackZhihui Xie, Tong Yu, Canzhe Zhao, Shuai LiSIGIR 2021 · 40 citations
- Variance-aware Regret Bounds for Stochastic Contextual Dueling BanditsQiwei Di, Tao Jin, Yue Wu, Heyang Zhao et al.ICLR 2024 · 21 citations
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet et al.ICLR 2025
- Efficient Explorative Key-Term Selection Strategies for Conversational Contextual BanditsZhiyong Wang, Xutong Liu, Shuai Li, John C. S. LuiAAAI 2023 · 18 citations
