Intrinsic Reward Driven Imitation Learning via Generative Model
Xingrui Yu, Yueming Lyu, Ivor W. Tsang
摘要
Imitation learning in a high-dimensional environment is challenging. Most inverse reinforcement learning (IRL) methods fail to outperform the demonstrator in such a high-dimensional environment, e.g., Atari domain. To address this challenge, we propose a novel reward learning module to generate intrinsic reward signals via a generative model. Our generative method can perform better forward state transition and backward action encoding, which improves the module's dynamics modeling ability in the environment. Thus, our module provides the imitation agent both the intrinsic intention of the demonstrator and a better exploration ability, which is critical for the agent to outperform the demonstrator. Empirical results show that our method outperforms state-of-the-art IRL methods on multiple Atari games, even with one-life demonstration. Remarkably, our method achieves performance that is up to 5 times the performance of the demonstration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Imitation by Predicting ObservationsAndrew Jaegle, Yury Sulsky, Arun Ahuja, Jake Bruce 等ICML 2021 · 被引用 16 次
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 被引用 13 次
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 被引用 7 次
- Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation LearningLiangyu Huo, Zulin Wang, Mai XuAAAI 2023
- Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive TransformerHao Luo, Zongqing LuICLR 2025
相关 Paper
- Receding Horizon Inverse Reinforcement LearningYiqing Xu, Wei Gao, David HsuNeurIPS 2022 · 被引用 17 次
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish 等ICLR 2025
- Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy UpdatesAnish Abhijit Diwan, Davide Tateo, Christopher Mower, Haitham Bou Ammar 等ICML 2026
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Reward-free World Models for Online Imitation LearningShangzhe Li, Zhiao Huang, Hao SuICML 2025
