Demonstration-Regularized RL
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Alexey Naumov, Pierre Perrault, Michal Valko, Pierre Ménard
Abstract
Incorporating expert demonstrations has empirically helped to improve the sample efficiency of reinforcement learning (RL). This paper quantifies theoretically to what extent this extra information reduces RL's sample complexity. In particular, we study the demonstration-regularized reinforcement learning that leverages the expert demonstrations by KL-regularization for a policy learned by behavior cloning. Our findings reveal that using expert demonstrations enables the identification of an optimal policy at a sample complexity of order in finite and in linear Markov decision processes, where is the target precision, the horizon, the number of action, the number of states in the finite case and the dimension of the feature space in the linear case. As a by-product, we provide tight convergence guarantees for the behaviour cloning procedure under general assumptions on the policy classes. Additionally, we establish that demonstration-regularized methods are provably efficient for reinforcement learning from human feedback (RLHF). In this respect, we provide theoretical evidence showing the benefits of KL-regularization for RLHF in tabular and linear MDPs. Interestingly, we avoid pessimism injection by employing computationally feasible regularization to handle reward estimation uncertainty, thus setting our approach apart from the prior works.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Improved Stochastic Optimization of LogSumExpEgor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel DvurechenskiiICML 2026 · 3 citations
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and VideogamesEv Zisselman, Mirco Mutti, Shelly Francis-Meretzki, Elisei Shafer et al.NeurIPS 2025 · 1 citation
- APC-RL: Exceeding data-driven behavior priors with adaptive policy compositionFinn Rietz, Pedro Zuidberg Dos Martires, Johannes A. StorkICLR 2026
- Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation LearningTian Xu, Zexuan Chen, Zhilong Zhang, Yi-Chen Li et al.ICML 2026
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 271 citations
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
Related papers
- Sharp Analysis for KL-Regularized Contextual Bandits and RLHFHeyang Zhao, Chenlu Ye, Quanquan Gu, Tong ZhangNeurIPS 2025 · 33 citations
- On Pathologies in KL-Regularized Reinforcement Learning from Expert DemonstrationsTim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal et al.NeurIPS 2021 · 33 citations
- Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage PerspectiveJiawei Huang, Bingcong Li, Christoph Dann, Niao HeICML 2025
- Data augmentation for efficient learning from parametric expertsAlexandre Galashov, Joshua Scott Merel, Nicolas HeessNeurIPS 2022 · 9 citations
- Logarithmic Regret for Online KL-Regularized Reinforcement LearningHeyang Zhao, Chenlu Ye, Wei Xiong, Quanquan Gu et al.ICML 2025
