Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
Qiwei Di, Jiafan He, Quanquan Gu
摘要
Learning from human feedback plays an important role in aligning generative models, such as large language models (LLM). However, the effectiveness of this approach can be influenced by adversaries, who may intentionally provide misleading preferences to manipulate the output in an undesirable or harmful direction. To tackle this challenge, we study a specific model within this problem domain-contextual dueling bandits with adversarial feedback, where the true preference label can be flipped by an adversary. We propose an algorithm namely robust contextual dueling bandits (RCDB), which is based on uncertaintyweighted maximum likelihood estimation. Our algorithm achieves an O(d √ T /κ + dC/κ) regret bound, where T is the number of rounds, d is the dimension of the context, κ is the lower bound of the derivative of the link function, and 0 ≤ C ≤ T is the total number of adversarial feedback. We also prove a lower bound to show that our regret bound is nearly optimal, both in scenarios with and without (C = 0) adversarial feedback. Our work is the first to achieve nearly minimax optimal regret for dueling bandits in the presence of adversarial preference feedback. Additionally, for the sigmoid link function, we develop a novel algorithm that takes into account the effect of local derivatives into maximum likelihood estimation (MLE) analysis through a refined method for estimating the link function's derivative. This method helps us to eliminate the κ dependence in the leading term with respect to T , which reduces the exponential dependence on the parameter radius B to a polynomial dependence. We conduct experiments to evaluate our proposed algorithm RCDB against various types of adversar-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple OptionsJoongkyu Lee, Seouh-won Yi, Min-hwan OhNeurIPS 2025 · 被引用 3 次
- Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial CorruptionsYoungmin OhICML 2026
- Learning from Imperfect Human Feedback: A Tale from Corruption-Robust DuelingYuwei Cheng, Fan Yao, Xuefeng Liu, Haifeng XuICLR 2025
- Best-of-three-worlds Analysis for Dueling Bandits with Borda WinnerZirui Hu, Tingyu Zhang, Fang KongICLR 2026
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Improved Optimistic Algorithms for Logistic BanditsLouis Faury, Marc Abeille, Clément Calauzènes, Olivier FercoqICML 2020 · 被引用 127 次
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 被引用 66 次
相关 Paper
- Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression OraclesAadirupa Saha, Robert E. SchapireNeurIPS 2025 · 被引用 3 次
- Variance-aware Regret Bounds for Stochastic Contextual Dueling BanditsQiwei Di, Tao Jin, Yue Wu, Heyang Zhao 等ICLR 2024 · 被引用 21 次
- Tackling Biased Evaluators in Dueling BanditsMing Tang, Yuxuan Zhou, Chao HuangNeurIPS 2025
- Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity ModelsViktor Bengs, Aadirupa Saha, Eyke HüllermeierICML 2022 · 被引用 32 次
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet 等ICLR 2025
