Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation Learning
Liangyu Huo, Zulin Wang, Mai Xu
Abstract
Imitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of the demonstrations and thus cannot learn policies to surpass the demonstrators. Some methods have been investigated to obtain better-than-demonstration (BD) performance with inner human feedback or preference labels. In this paper, we propose a method to learn rewards from suboptimal demonstrations via a weighted preference learning technique (LERP). Specifically, we first formulate the suboptimality of demonstrations as the inaccurate estimation of rewards. The inaccuracy is modeled with a reward noise random variable following the Gumbel distribution. Moreover, we derive an upper bound of the expected return with different noise coefficients and propose a theorem to surpass the demonstrations. Unlike existing literature, our analysis does not depend on the linear reward constraint. Consequently, we develop a BD model with a weighted preference learning technique. Experimental results on continuous control and high-dimensional discrete control tasks show the superiority of our LERP method over other state-of-the-art BD methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04f0c8a2-6548-4df0-82eb-aa85e82b8106Cited by top-tier papers3
- Hierarchical Policy Learning via Spectral DecompositionShuxin Cao, Liquan Wang, Walker Byrnes, Yiye Chen et al.ICML 2026
- PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsQiang Liu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen et al.ICLR 2025
- Improving Reward Model Generalization from Adversarial Process Enhanced PreferencesZhilong Zhang, Tian Xu, Xinghao Du, Xingchen Cao et al.ICML 2025
Builds on2
Related papers
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 57 citations
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 11 citations
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu et al.ICML 2024 · 6 citations
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 25 citations
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng et al.AAAI 2025
