Learning from Imperfect Human Feedback: A Tale from Corruption-Robust Dueling
Yuwei Cheng, Fan Yao, Xuefeng Liu, Haifeng Xu
摘要
This paper studies Learning from Imperfect Human Feedback (LIHF), addressing the potential irrationality or imperfect perception when learning from comparative human feedback. Building on evidences that human's imperfection decays over time (i.e., humans learn to improve), we cast this problem as a concave-utility continuous-action dueling bandit but under a restricted form of corruption: i.e., the corruption scale is decaying over time as t ρ-1 for some "imperfection rate" ρ ∈ [0, 1]. With T as the total number of iterations, we establish a regret lower bound of Ω(max √ T , T ρ ) for LIHF, even when ρ is known. For the same setting, we develop the Robustified Stochastic Mirror Descent for Imperfect Dueling (RoSMID) algorithm, which achieves nearly optimal regret Õ(max √ T , T ρ ). Core to our analysis is a novel framework for analyzing gradient-based algorithms for dueling bandit under corruption, and we demonstrate its general applicability by showing how this framework can be easily applied to obtain corruption-robust guarantees for other popular gradient-based dueling bandit algorithms. Our theoretical results are validated by extensive experiments. † Equal contribution. ‡ Work done while visiting University of Chicago.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 被引用 66 次
- Optimal Algorithms for Stochastic Contextual Preference BanditsAadirupa SahaNeurIPS 2021 · 被引用 64 次
- Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative PreferencesAadirupa Saha, Pierre GaillardICML 2022 · 被引用 30 次
相关 Paper
- Adversarial Combinatorial Semi-bandits with Graph FeedbackYuxiao WenICML 2025
- Fusing Reward and Dueling Feedback in Stochastic BanditsXuchuang Wang, Qirun Zeng, Jinhang Zuo, Xutong Liu 等ICML 2025
- Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression OraclesAadirupa Saha, Robert E. SchapireNeurIPS 2025 · 被引用 3 次
- Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial FeedbackQiwei Di, Jiafan He, Quanquan GuICML 2025
- On Optimal Robustness to Adversarial Corruption in Online Decision ProblemsShinji ItoNeurIPS 2021 · 被引用 28 次
