Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence Rate
Yufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran Wang
摘要
Generative adversarial imitation learning (GAIL) demonstrates tremendous success in practice, especially when combined with neural networks. Different from reinforcement learning, GAIL learns both policy and reward function from expert (human) demonstration. Despite its empirical success, it remains unclear whether GAIL with neural networks converges to the globally optimal solution. The major difficulty comes from the nonconvex-nonconcave minimax optimization structure. To bridge the gap between practice and theory, we analyze a gradient-based algorithm with alternating updates and establish its sublinear convergence to the globally optimal solution. To the best of our knowledge, our analysis establishes the global optimality and convergence rate of GAIL with neural networks for the first time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk 等NeurIPS 2022 · 被引用 27 次
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo 等ICML 2024 · 被引用 17 次
- C-GAIL: Stabilizing Generative Adversarial Imitation Learning with Control TheoryTianjiao Luo, Tim Pearce, Huayu Chen, Jianfei Chen 等NeurIPS 2024 · 被引用 13 次
- Efficient Performance Bounds for Primal-Dual Reinforcement Learning from DemonstrationsAngeliki Kamoutsi, Goran Banjac, John LygerosICML 2021 · 被引用 9 次
- Provably and Practically Efficient Adversarial Imitation Learning with General Function ApproximationTian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun 等NeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper2
相关 Paper
- Single-Timescale Actor-Critic Provably Finds Globally Optimal PolicyZuyue Fu, Zhuoran Yang, Zhaoran WangICLR 2021 · 被引用 52 次
- Learning from Demonstration: Provably Efficient Adversarial Policy Imitation with Linear Function ApproximationZhihan Liu, Yufeng Zhang, Zuyue Fu, Zhuoran Yang 等ICML 2022 · 被引用 17 次
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo 等ICML 2025
- PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsQiang Liu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 等ICLR 2025
- Imitation Learning for Human Pose PredictionBorui Wang, Ehsan Adeli, Hsu-Kuang Chiu, De-An Huang 等ICCV 2019 · 被引用 110 次
