Adversarial Imitation Learning via Boosting
Jonathan D. Chang, Dhruv Sreenivas, Yingbing Huang, Kianté Brantley, Wen Sun
摘要
Adversarial imitation learning (AIL) has stood out as a dominant framework across various imitation learning (IL) applications, with Discriminator Actor Critic (DAC) (Kostrikov et al., 2019) demonstrating the effectiveness of off-policy learning algorithms in improving sample efficiency and scalability to higher-dimensional observations. Despite DAC's empirical success, the original AIL objective is on-policy and DAC's ad-hoc application of off-policy training does not guarantee successful imitation (Kostrikov et al., 2019; 2020). Follow-up work such as ValueDICE (Kostrikov et al., 2020) tackles this issue by deriving a fully off-policy AIL objective. Instead in this work, we develop a novel and principled AIL algorithm via the framework of boosting. Like boosting, our new algorithm, AILBoost, maintains an ensemble of properly weighted weak learners (i.e., policies) and trains a discriminator that witnesses the maximum discrepancy between the distributions of the ensemble and the expert policy. We maintain a weighted replay buffer to represent the state-action distribution induced by the ensemble, allowing us to train discriminators using the entire data collected so far. In the weighted replay buffer, the contribution of the data from older policies are properly discounted with the weight computed based on the boosting framework. Empirically, we evaluate our algorithm on both controller state-based and pixel-based environments from the DeepMind Control Suite. AILBoost outperforms DAC on both types of environments, demonstrating the benefit of properly weighting replay buffer data for off-policy training. On state-based environments, AILBoost outperforms ValueDICE and IQ-Learn (Garg et al., 2021) , achieving competitive performance with as little as one expert trajectory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 被引用 1 次
- Noise-Guided Transport: Imitation Learning from Random PriorsLionel Blondé, Joao A. Candido Ramos, Alexandros KalousisICML 2026
- Diffusing States and Matching Scores: A New Framework for Imitation LearningRunzhe Wu, Yiding Chen, Gokul Swamy, Kianté Brantley 等ICLR 2025
- IL-SOAR : Imitation Learning with Soft Optimistic Actor cRiticStefano Viel, Luca Viano, Volkan CevherICML 2025
它引用的顶会 Paper13
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
相关 Paper
- Policy Contrastive Imitation LearningJialei Huang, Zhao-Heng Yin, Yingdong Hu, Yang GaoICML 2023 · 被引用 4 次
- Deterministic and Discriminative Imitation (D2-Imitation): Revisiting Adversarial Imitation for Sample EfficiencyMingfei Sun, Sam Devlin, Katja Hofmann, Shimon WhitesonAAAI 2022 · 被引用 7 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- Adversarial Soft Advantage Fitting: Imitation Learning without Policy OptimizationPaul Barde, Julien Roy, Wonseok Jeon, Joelle Pineau 等NeurIPS 2020 · 被引用 30 次
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
