Intrinsic Reward Driven Imitation Learning via Generative Model
Xingrui Yu, Yueming Lyu, Ivor W. Tsang
Abstract
Imitation learning in a high-dimensional environment is challenging. Most inverse reinforcement learning (IRL) methods fail to outperform the demonstrator in such a high-dimensional environment, e.g., Atari domain. To address this challenge, we propose a novel reward learning module to generate intrinsic reward signals via a generative model. Our generative method can perform better forward state transition and backward action encoding, which improves the module's dynamics modeling ability in the environment. Thus, our module provides the imitation agent both the intrinsic intention of the demonstrator and a better exploration ability, which is critical for the agent to outperform the demonstrator. Empirical results show that our method outperforms state-of-the-art IRL methods on multiple Atari games, even with one-life demonstration. Remarkably, our method achieves performance that is up to 5 times the performance of the demonstration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e28050d2-ea0d-40aa-99d3-a8e9b9edabb2Cited by top-tier papers9
- Imitation by Predicting ObservationsAndrew Jaegle, Yury Sulsky, Arun Ahuja, Jake Bruce et al.ICML 2021 · 16 citations
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 13 citations
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 7 citations
- Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation LearningLiangyu Huo, Zulin Wang, Mai XuAAAI 2023
- Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive TransformerHao Luo, Zongqing LuICLR 2025
Related papers
- Receding Horizon Inverse Reinforcement LearningYiqing Xu, Wei Gao, David HsuNeurIPS 2022 · 17 citations
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish et al.ICLR 2025
- Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy UpdatesAnish Abhijit Diwan, Davide Tateo, Christopher Mower, Haitham Bou Ammar et al.ICML 2026
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Reward-free World Models for Online Imitation LearningShangzhe Li, Zhiao Huang, Hao SuICML 2025
