On Pathologies in KL-Regularized Reinforcement Learning from Expert Demonstrations
Tim G. J. Rudner, Cong Lu, Michael A. Osborne, Yarin Gal, Yee Whye Teh
Abstract
KL-regularized reinforcement learning from expert demonstrations has proved successful in improving the sample efficiency of deep reinforcement learning algorithms, allowing them to be applied to challenging physical real-world tasks. However, we show that KL-regularized reinforcement learning with behavioral reference policies derived from expert demonstrations can suffer from pathological training dynamics that can lead to slow, unstable, and suboptimal online learning. We show empirically that the pathology occurs for commonly chosen behavioral policy classes and demonstrate its impact on sample efficiency and online policy performance. Finally, we show that the pathology can be remedied by non-parametric behavioral reference policies and that this allows KL-regularized reinforcement learning to significantly outperform state-of-the-art approaches on a variety of challenging locomotion and dexterous hand manipulation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Tractable Function-Space Variational Inference in Bayesian Neural NetworksTim G. J. Rudner, Zonghao Chen, Yee Whye Teh, Yarin GalNeurIPS 2022 · 70 citations
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer et al.ICML 2022 · 69 citations
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentPhilip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. RobertsICML 2021 · 55 citations
- Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy MatchingYecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, Osbert BastaniICML 2022 · 49 citations
Builds on7
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki et al.ICLR 2020 · 299 citations
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentPhilip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. RobertsICML 2021 · 55 citations
Related papers
- Demonstration-Regularized RLDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.ICLR 2024 · 5 citations
- Sharp Analysis for KL-Regularized Contextual Bandits and RLHFHeyang Zhao, Chenlu Ye, Quanquan Gu, Tong ZhangNeurIPS 2025 · 33 citations
- Data augmentation for efficient learning from parametric expertsAlexandre Galashov, Joshua Scott Merel, Nicolas HeessNeurIPS 2022 · 9 citations
- Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement LearningChen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao et al.ICML 2025
- AMOR: Adaptive Character Control through Multi-Objective Reinforcement LearningLucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller et al.SIGGRAPH 2025 · 4 citations
