EvIL: Evolution Strategies for Generalisable Imitation Learning
Silvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh, Jakob Nicolaus Foerster
摘要
Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g. demonstrations collected in simulation but deployment in the real world). Compared to policy-centric approaches to IL like behavioural cloning, reward-centric approaches like inverse reinforcement learning (IRL) often better replicate expert behaviour in new environments. This transfer is usually performed by optimising the recovered reward under the dynamics of the target environment. However, (a) we find that modern deep IL algorithms frequently recover rewards which induce policies far weaker than the expert, even in the same environment the demonstrations were collected in. Furthermore, (b) these rewards are often quite poorly shaped, necessitating extensive environment interaction to optimise effectively. We provide simple and scalable fixes to both of these concerns. For (a), we find that reward model ensembles combined with a slightly different training objective significantly improves re-training and transfer performance. For (b), we propose a novel evolution-strategies based method (EvIL) to optimise for a reward-shaping term that speeds up re-training in the target environment, closing a gap left open by the classical theory of IRL. On a suite of continuous control tasks, we are able to re-train policies in target (and source) environments more interaction-efficiently than prior work. Our code is open-sourced at https: //github.com/SilviaSapora/evil .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On Discovering Algorithms for Adversarial Imitation LearningShashank Reddy Chirra, Jayden Teoh, Praveen Paruchuri, Pradeep VarakanthamICLR 2026 · 被引用 1 次
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement LearningSilvia Sapora, R. Devon Hjelm, Omar Attia, Alexander Toshev 等ICLR 2026
它引用的顶会 Paper12
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 53 次
- Bridging RL Theory and Practice with the Effective HorizonCassidy Laidlaw, Stuart J. Russell, Anca D. DraganNeurIPS 2023 · 被引用 42 次
- Hybrid Inverse Reinforcement LearningJuntao Ren, Gokul Swamy, Steven Wu, Drew Bagnell 等ICML 2024 · 被引用 33 次
相关 Paper
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun 等AAAI 2024 · 被引用 8 次
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 被引用 26 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Provably Efficient Learning of Transferable RewardsAlberto Maria Metelli, Giorgia Ramponi, Alessandro Concetti, Marcello RestelliICML 2021 · 被引用 36 次
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira 等ICLR 2023 · 被引用 1 次
