Bandits with Preference Feedback: A Stackelberg Game Perspective
Barna Pásztor, Parnian Kassraie, Andreas Krause
摘要
Bandits with preference feedback present a powerful tool for optimizing unknown target functions when only pairwise comparisons are allowed instead of direct value queries. This model allows for incorporating human feedback into online inference and optimization and has been employed in systems for fine-tuning large language models. The problem is well understood in simplified settings with linear target functions or over finite small domains that limit practical interest. Taking the next step, we consider infinite domains and nonlinear (kernelized) rewards. In this setting, selecting a pair of actions is quite challenging and requires balancing exploration and exploitation at two levels: within the pair, and along the iterations of the algorithm. We propose MAXMINLCB, which emulates this trade-off as a zero-sum Stackelberg game, and chooses action pairs that are informative and yield favorable rewards. MAXMINLCB consistently outperforms existing algorithms and satisfies an anytime-valid rate-optimal regret guarantee. This is due to our novel preference-based confidence sequences for kernelized logistic estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Optimal Design for Human Preference ElicitationSubhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Deshmukh 等NeurIPS 2024 · 被引用 20 次
- ActiveUltraFeedback: Efficient Preference Data Generation using Active LearningDavit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich 等ICML 2026
- Adversarial Policy Optimization for Offline Preference-based Reinforcement LearningHyungkyu Kang, Min-hwan OhICLR 2025
- Bayesian Optimization from Human Feedback: Near-Optimal Regret BoundsAya Kayal, Sattar Vakili, Laura Toni, Da-shan Shiu 等ICML 2025
- Efficiently Learning at Test-Time: Active Fine-Tuning of LLMsJonas Hübotter, Sascha Bongni, Ido Hakimi, Andreas KrauseICLR 2025
它引用的顶会 Paper9
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar 等ICML 2024 · 被引用 212 次
- A framework for bilevel optimization that enables stochastic and global variance reduction algorithmsMathieu Dagréou, Pierre Ablin, Samuel Vaiter, Thomas MoreauNeurIPS 2022 · 被引用 149 次
- Improved Optimistic Algorithms for Logistic BanditsLouis Faury, Marc Abeille, Clément Calauzènes, Olivier FercoqICML 2020 · 被引用 127 次
- Optimal Algorithms for Stochastic Contextual Preference BanditsAadirupa SahaNeurIPS 2021 · 被引用 64 次
相关 Paper
- Provable Reward-Agnostic Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. LeeICLR 2024 · 被引用 16 次
- Bootstrapping LLMs via Preference-Based Policy OptimizationChen JiaAAAI 2026
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential GameBarna Pásztor, Thomas Kleine Buening, Andreas KrauseICLR 2026 · 被引用 9 次
- Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human FeedbackShengbo Wang, Hong Sun, Ke LiAAAI 2026
- Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression OraclesAadirupa Saha, Robert E. SchapireNeurIPS 2025 · 被引用 3 次
