Bandits with Preference Feedback: A Stackelberg Game Perspective
Barna Pásztor, Parnian Kassraie, Andreas Krause
Abstract
Bandits with preference feedback present a powerful tool for optimizing unknown target functions when only pairwise comparisons are allowed instead of direct value queries. This model allows for incorporating human feedback into online inference and optimization and has been employed in systems for fine-tuning large language models. The problem is well understood in simplified settings with linear target functions or over finite small domains that limit practical interest. Taking the next step, we consider infinite domains and nonlinear (kernelized) rewards. In this setting, selecting a pair of actions is quite challenging and requires balancing exploration and exploitation at two levels: within the pair, and along the iterations of the algorithm. We propose MAXMINLCB, which emulates this trade-off as a zero-sum Stackelberg game, and chooses action pairs that are informative and yield favorable rewards. MAXMINLCB consistently outperforms existing algorithms and satisfies an anytime-valid rate-optimal regret guarantee. This is due to our novel preference-based confidence sequences for kernelized logistic estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e9489ca-9754-4360-b88a-25ecc045a2e2Cited by top-tier papers5
- Optimal Design for Human Preference ElicitationSubhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Deshmukh et al.NeurIPS 2024 · 20 citations
- ActiveUltraFeedback: Efficient Preference Data Generation using Active LearningDavit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich et al.ICML 2026
- Adversarial Policy Optimization for Offline Preference-based Reinforcement LearningHyungkyu Kang, Min-hwan OhICLR 2025
- Bayesian Optimization from Human Feedback: Near-Optimal Regret BoundsAya Kayal, Sattar Vakili, Laura Toni, Da-shan Shiu et al.ICML 2025
- Efficiently Learning at Test-Time: Active Fine-Tuning of LLMsJonas Hübotter, Sascha Bongni, Ido Hakimi, Andreas KrauseICLR 2025
Builds on9
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar et al.ICML 2024 · 212 citations
- A framework for bilevel optimization that enables stochastic and global variance reduction algorithmsMathieu Dagréou, Pierre Ablin, Samuel Vaiter, Thomas MoreauNeurIPS 2022 · 149 citations
- Improved Optimistic Algorithms for Logistic BanditsLouis Faury, Marc Abeille, Clément Calauzènes, Olivier FercoqICML 2020 · 127 citations
- Optimal Algorithms for Stochastic Contextual Preference BanditsAadirupa SahaNeurIPS 2021 · 64 citations
Related papers
- Provable Reward-Agnostic Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. LeeICLR 2024 · 16 citations
- Bootstrapping LLMs via Preference-Based Policy OptimizationChen JiaAAAI 2026
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential GameBarna Pásztor, Thomas Kleine Buening, Andreas KrauseICLR 2026 · 9 citations
- Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human FeedbackShengbo Wang, Hong Sun, Ke LiAAAI 2026
- Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression OraclesAadirupa Saha, Robert E. SchapireNeurIPS 2025 · 3 citations
