EvIL: Evolution Strategies for Generalisable Imitation Learning
Silvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh, Jakob Nicolaus Foerster
Abstract
Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g. demonstrations collected in simulation but deployment in the real world). Compared to policy-centric approaches to IL like behavioural cloning, reward-centric approaches like inverse reinforcement learning (IRL) often better replicate expert behaviour in new environments. This transfer is usually performed by optimising the recovered reward under the dynamics of the target environment. However, (a) we find that modern deep IL algorithms frequently recover rewards which induce policies far weaker than the expert, even in the same environment the demonstrations were collected in. Furthermore, (b) these rewards are often quite poorly shaped, necessitating extensive environment interaction to optimise effectively. We provide simple and scalable fixes to both of these concerns. For (a), we find that reward model ensembles combined with a slightly different training objective significantly improves re-training and transfer performance. For (b), we propose a novel evolution-strategies based method (EvIL) to optimise for a reward-shaping term that speeds up re-training in the target environment, closing a gap left open by the classical theory of IRL. On a suite of continuous control tasks, we are able to re-train policies in target (and source) environments more interaction-efficiently than prior work. Our code is open-sourced at https: //github.com/SilviaSapora/evil .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e51b3c6c-8755-4c1b-924e-56e7844d4b24Cited by top-tier papers2
- On Discovering Algorithms for Adversarial Imitation LearningShashank Reddy Chirra, Jayden Teoh, Praveen Paruchuri, Pradeep VarakanthamICLR 2026 · 1 citation
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement LearningSilvia Sapora, R. Devon Hjelm, Omar Attia, Alexander Toshev et al.ICLR 2026
Builds on12
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Bridging RL Theory and Practice with the Effective HorizonCassidy Laidlaw, Stuart J. Russell, Anca D. DraganNeurIPS 2023 · 42 citations
- Hybrid Inverse Reinforcement LearningJuntao Ren, Gokul Swamy, Steven Wu, Drew Bagnell et al.ICML 2024 · 33 citations
Related papers
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun et al.AAAI 2024 · 8 citations
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 26 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Provably Efficient Learning of Transferable RewardsAlberto Maria Metelli, Giorgia Ramponi, Alessandro Concetti, Marcello RestelliICML 2021 · 36 citations
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira et al.ICLR 2023 · 1 citation
