Reinforcement Learning from Imperfect Demonstrations under Soft Expert Guidance
Mingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Chao Yang, Bin Fang, Huaping Liu
Abstract
In this paper, we study Reinforcement Learning from Demonstrations (RLfD) that improves the exploration efficiency of Reinforcement Learning (RL) by providing expert demonstrations. Most of existing RLfD methods require demonstrations to be perfect and sufficient, which yet is unrealistic to meet in practice. To work on imperfect demonstrations, we first define an imperfect expert setting for RLfD in a formal way, and then point out that previous methods suffer from two issues in terms of optimality and convergence, respectively. Upon the theoretical findings we have derived, we tackle these two issues by regarding the expert guidance as a soft constraint on regulating the policy exploration of the agent, which eventually leads to a constrained optimization problem. We further demonstrate that such problem is able to be addressed efficiently by performing a local linear search on its dual form. Considerable empirical evaluations on a comprehensive collection of benchmarks indicate our method attains consistent improvement over other RLfD counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 795f86e0-5e22-4fbe-9545-634ffa8b18eaCited by top-tier papers4
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador et al.NeurIPS 2021 · 53 citations
- Align-RUDDER: Learning From Few Demonstrations by Reward RedistributionVihang Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer et al.ICML 2022 · 46 citations
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 14 citations
- Iterative Regularized Policy Optimization with Imperfect DemonstrationsXudong Gong, Dawei Feng, Kele Xu, Yuanzhao Zhai et al.ICML 2024 · 5 citations
Related papers
- Hybrid Policy Optimization from Imperfect DemonstrationsHanlin Yang, Chao Yu, Peng Sun, Siji ChenNeurIPS 2023 · 14 citations
- Learning and Repair of Deep Reinforcement Learning Policies from Fuzz-Testing DataMartin Tappler, Andrea Pferscher, Bernhard K. Aichernig, Bettina KönighoferICSE 2024 · 6 citations
- RLfOLD: Reinforcement Learning from Online Demonstrations in Urban Autonomous DrivingDaniel Coelho, Miguel Oliveira, Vitor SantosAAAI 2024 · 14 citations
- A Few Expert Queries Suffices for Sample-Efficient RL with Resets and Linear Value ApproximationPhilip Amortila, Nan Jiang, Dhruv Madeka, Dean P. FosterNeurIPS 2022 · 6 citations
- Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited DemonstrationsHaowen Sun, Liqi Huang, Mingyang Li, Sihua Ren et al.ICML 2026
