Deterministic and Discriminative Imitation (D2-Imitation): Revisiting Adversarial Imitation for Sample Efficiency
Mingfei Sun, Sam Devlin, Katja Hofmann, Shimon Whiteson
Abstract
Sample efficiency is crucial for imitation learning methods to be applicable in real-world applications. Many studies improve sample efficiency by extending adversarial imitation to be off-policy regardless of the fact that these off-policy extensions could either change the original objective or involve complicated optimization. We revisit the foundation of adversarial imitation and propose an off-policy sample efficient approach that requires no adversarial training or min-max optimization. Our formulation capitalizes on two key insights: (1) the similarity between the Bellman equation and the stationary state-action distribution equation allows us to derive a novel temporal difference (TD) learning approach; and (2) the use of a deterministic policy simplifies the TD learning. Combined, these insights yield a practical algorithm, Deterministic and Discriminative Imitation (D2-Imitation), which oper- ates by first partitioning samples into two replay buffers and then learning a deterministic policy via off-policy reinforcement learning. Our empirical results show that D2-Imitation is effective in achieving good sample efficiency, outperforming several off-policy extension approaches of adversarial imitation on many control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f339ff3-5c01-43ce-942a-7ef3ebfce8c8Cited by top-tier papers1
Ask how each one uses itBuilds on3
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 41 citations
Related papers
- Adversarial Soft Advantage Fitting: Imitation Learning without Policy OptimizationPaul Barde, Julien Roy, Wonseok Jeon, Joelle Pineau et al.NeurIPS 2020 · 30 citations
- Adversarial Imitation Learning via BoostingJonathan D. Chang, Dhruv Sreenivas, Yingbing Huang, Kianté Brantley et al.ICLR 2024 · 6 citations
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 22 citations
- Off-Policy Imitation Learning from ObservationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouNeurIPS 2020 · 102 citations
- Planning for Sample Efficient Imitation LearningZhao-Heng Yin, Weirui Ye, Qifeng Chen, Yang GaoNeurIPS 2022 · 32 citations
