C-GAIL: Stabilizing Generative Adversarial Imitation Learning with Control Theory
Tianjiao Luo, Tim Pearce, Huayu Chen, Jianfei Chen, Jun Zhu
摘要
Generative Adversarial Imitation Learning (GAIL) trains a generative policy to mimic a demonstrator. It uses on-policy Reinforcement Learning (RL) to optimize a reward signal derived from a GAN-like discriminator. A major drawback of GAIL is its training instability - it inherits the complex training dynamics of GANs, and the distribution shift introduced by RL. This can cause oscillations during training, harming its sample efficiency and final policy performance. Recent work has shown that control theory can help with the convergence of a GAN's training. This paper extends this line of work, conducting a control-theoretic analysis of GAIL and deriving a novel controller that not only pushes GAIL to the desired equilibrium but also achieves asymptotic stability in a 'one-step' setting. Based on this, we propose a practical algorithm 'Controlled-GAIL' (C-GAIL). On MuJoCo tasks, our controlled variant is able to speed up the rate of convergence, reduce the range of oscillation and match the expert's distribution more closely both for vanilla GAIL and GAIL-DAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 被引用 666 次
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- On Computation and Generalization of Generative Adversarial Imitation LearningMinshuo Chen, Yizhou Wang, Tianyi Liu, Zhuoran Yang 等ICLR 2020 · 被引用 42 次
- Understanding and Stabilizing GANs' Training Dynamics Using Control TheoryKun Xu, Chongxuan Li, Jun Zhu, Bo ZhangICML 2020 · 被引用 31 次
- Minimax Optimal Online Imitation Learning via Replay EstimationGokul Swamy, Nived Rajaraman, Matthew Peng, Sanjiban Choudhury 等NeurIPS 2022 · 被引用 27 次
相关 Paper
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
- Exploring Gradient Explosion in Generative Adversarial Imitation Learning: A Probabilistic PerspectiveWanying Wang, Yichen Zhu, Yirui Zhou, Chaomin Shen 等AAAI 2024 · 被引用 13 次
- Diffusion-Reward Adversarial Imitation LearningChun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Yu-Chiang Frank Wang 等NeurIPS 2024 · 被引用 28 次
- Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence RateYufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran WangICML 2020 · 被引用 12 次
- f-GAIL: Learning f-Divergence for Generative Adversarial Imitation LearningXin Zhang, Yanhua Li, Ziming Zhang, Zhi-Li ZhangNeurIPS 2020 · 被引用 40 次
