Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence Rate
Yufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran Wang
Abstract
Generative adversarial imitation learning (GAIL) demonstrates tremendous success in practice, especially when combined with neural networks. Different from reinforcement learning, GAIL learns both policy and reward function from expert (human) demonstration. Despite its empirical success, it remains unclear whether GAIL with neural networks converges to the globally optimal solution. The major difficulty comes from the nonconvex-nonconcave minimax optimization structure. To bridge the gap between practice and theory, we analyze a gradient-based algorithm with alternating updates and establish its sublinear convergence to the globally optimal solution. To the best of our knowledge, our analysis establishes the global optimality and convergence rate of GAIL with neural networks for the first time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk et al.NeurIPS 2022 · 27 citations
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo et al.ICML 2024 · 17 citations
- C-GAIL: Stabilizing Generative Adversarial Imitation Learning with Control TheoryTianjiao Luo, Tim Pearce, Huayu Chen, Jianfei Chen et al.NeurIPS 2024 · 13 citations
- Efficient Performance Bounds for Primal-Dual Reinforcement Learning from DemonstrationsAngeliki Kamoutsi, Goran Banjac, John LygerosICML 2021 · 9 citations
- Provably and Practically Efficient Adversarial Imitation Learning with General Function ApproximationTian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun et al.NeurIPS 2024 · 8 citations
Builds on2
Related papers
- Single-Timescale Actor-Critic Provably Finds Globally Optimal PolicyZuyue Fu, Zhuoran Yang, Zhaoran WangICLR 2021 · 52 citations
- Learning from Demonstration: Provably Efficient Adversarial Policy Imitation with Linear Function ApproximationZhihan Liu, Yufeng Zhang, Zuyue Fu, Zhuoran Yang et al.ICML 2022 · 17 citations
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo et al.ICML 2025
- PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsQiang Liu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen et al.ICLR 2025
- Imitation Learning for Human Pose PredictionBorui Wang, Ehsan Adeli, Hsu-Kuang Chiu, De-An Huang et al.ICCV 2019 · 110 citations
