Provably Feedback-Efficient Reinforcement Learning via Active Reward Learning
Dingwen Kong, Lin Yang
Abstract
An appropriate reward function is of paramount importance in specifying a task in reinforcement learning (RL). Yet, it is known to be extremely challenging in practice to design a correct reward function for even simple tasks. Human-in-the-loop (HiL) RL allows humans to communicate complex goals to the RL agent by providing various types of feedback. However, despite achieving great empirical successes, HiL RL usually requires too much feedback from a human teacher and also suffers from insufficient theoretical understanding. In this paper, we focus on addressing this issue from a theoretical perspective, aiming to provide provably feedback-efficient algorithmic frameworks that take human-in-the-loop to specify rewards of given tasks. We provide an active-learning-based RL algorithm that first explores the environment without specifying a reward function and then asks a human teacher for only a few queries about the rewards of a task at some state-action pairs. After that, the algorithm guarantees to provide a nearly optimal policy for the task with high probability. We show that, even with the presence of random noise in the feedback, the algorithm only takes queries on the reward function to provide an -optimal policy for any . Here is the horizon of the RL environment, and specifies the complexity of the function class representing the reward function. In contrast, standard RL algorithms require to query the reward function for at least state-action pairs where depends on the complexity of the environmental transition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 188d5c05-b21c-4d32-82c8-593a9ceffcb6Cited by top-tier papers4
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- ReDit: Reward Dithering for Improved LLM Policy OptimizationChenxing Wei, Jiarui Yu, Ying He, Hande Dong et al.NeurIPS 2025 · 14 citations
- Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward InferenceQining Zhang, Lei YingICLR 2025
- Differentially Private Preference Data Synthesis for Large Language Model AlignmentFengyu Gao, Jing YangICML 2026
Builds on25
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
Related papers
- Efficient Meta Reinforcement Learning for Preference-based Fast AdaptationZhizhou Ren, Anji Liu, Yitao Liang, Jian Peng et al.NeurIPS 2022 · 11 citations
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek et al.NeurIPS 2021 · 27 citations
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang et al.ICLR 2024 · 12 citations
- Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function ApproximationXiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang et al.ICML 2022 · 90 citations
- DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human FeedbackXuening Feng, Zhaohui Jiang, Timo Kaufmann, Puchen Xu et al.AAAI 2025 · 7 citations
