Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation Learning
Liangyu Huo, Zulin Wang, Mai Xu
摘要
Imitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of the demonstrations and thus cannot learn policies to surpass the demonstrators. Some methods have been investigated to obtain better-than-demonstration (BD) performance with inner human feedback or preference labels. In this paper, we propose a method to learn rewards from suboptimal demonstrations via a weighted preference learning technique (LERP). Specifically, we first formulate the suboptimality of demonstrations as the inaccurate estimation of rewards. The inaccuracy is modeled with a reward noise random variable following the Gumbel distribution. Moreover, we derive an upper bound of the expected return with different noise coefficients and propose a theorem to surpass the demonstrations. Unlike existing literature, our analysis does not depend on the linear reward constraint. Consequently, we develop a BD model with a weighted preference learning technique. Experimental results on continuous control and high-dimensional discrete control tasks show the superiority of our LERP method over other state-of-the-art BD methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Hierarchical Policy Learning via Spectral DecompositionShuxin Cao, Liquan Wang, Walker Byrnes, Yiye Chen 等ICML 2026
- PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsQiang Liu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 等ICLR 2025
- Improving Reward Model Generalization from Adversarial Process Enhanced PreferencesZhilong Zhang, Tian Xu, Xinghao Du, Xingchen Cao 等ICML 2025
它引用的顶会 Paper2
相关 Paper
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 被引用 11 次
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu 等ICML 2024 · 被引用 6 次
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 被引用 25 次
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng 等AAAI 2025
