Contextual Bandits and Imitation Learning with Preference-Based Active Queries
Ayush Sekhari, Karthik Sridharan, Wen Sun, Runzhe Wu
Abstract
We consider the problem of contextual bandits and imitation learning, where the learner lacks direct knowledge of the executed action's reward. Instead, the learner can actively query an expert at each round to compare two actions and receive noisy preference feedback. The learner's objective is twofold: to minimize the regret associated with the executed actions, while simultaneously, minimizing the number of comparison queries made to the expert. In this paper, we assume that the learner has access to a function class that can represent the expert's preference model under appropriate link functions, and provide an algorithm that leverages an online regression oracle with respect to this function class for choosing its actions and deciding when to query. For the contextual bandit setting, our algorithm achieves a regret bound that combines the best of both worlds, scaling as O(min √ T , d/∆), where T represents the number of interactions, d represents the eluder dimension of the function class, and ∆ represents the minimum preference of the optimal action over any suboptimal action under all contexts. Our algorithm does not require the knowledge of ∆, and the obtained regret bound is comparable to what can be achieved in the standard contextual bandits setting where the learner observes reward signals at each round. Additionally, our algorithm makes only O(minT, d 2 /∆ 2 ) queries to the expert. We then extend our algorithm to the imitation learning setting, where the learning agent engages with an unknown environment in episodes of length H each, and provide similar guarantees for regret and query complexity. The regret bound for our imitation learning algorithm, which relies on preferencebased feedback, matches the prior results in interactive imitation learning (Ross and Bagnell, 2014) that require access to the expert's actions as well as reward signals. Furthermore, we show that our algorithm enjoys improved query complexity bounds. Interestingly, in some cases, our algorithm for imitation learning via preference-feedback can even learn to outperform the underlying expert thus highlighting a practical benefit of considering preference-based feedback in imitation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e94bde4-9cd4-4954-9456-08a7605dfaa4Cited by top-tier papers6
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 44 citations
- Optimal Design for Human Preference ElicitationSubhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Deshmukh et al.NeurIPS 2024 · 20 citations
- Provably Efficient Online RLHF with One-Pass Reward ModelingLong-Fei Li, Yu-Yang Qian, Peng Zhao, Zhi-Hua ZhouNeurIPS 2025 · 8 citations
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang et al.KDD 2024 · 6 citations
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental DesignAndreas Schlaginhaufen, Reda Ouhamma, Maryam KamgarpourNeurIPS 2025 · 4 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
Related papers
- Selective Sampling and Imitation Learning via Online RegressionAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 15 citations
- Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity ModelsViktor Bengs, Aadirupa Saha, Eyke HüllermeierICML 2022 · 32 citations
- Adapting to Misspecification in Contextual BanditsDylan J. Foster, Claudio Gentile, Mehryar Mohri, Julian ZimmertNeurIPS 2020 · 111 citations
- How Does Variance Shape the Regret in Contextual Bandits?Zeyu Jia, Jian Qian, Alexander Rakhlin, Chen-Yu WeiNeurIPS 2024 · 13 citations
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
