Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
Tian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun, Yang Yu
摘要
As a prominent category of imitation learning methods, adversarial imitation learning (AIL) has garnered significant practical success powered by neural network approximation. However, existing theoretical studies on AIL are primarily limited to simplified scenarios such as tabular and linear function approximation and involve complex algorithmic designs that hinder practical implementation, highlighting a gap between theory and practice. In this paper, we explore the theoretical underpinnings of online AIL with general function approximation. We introduce a new method called optimization-based AIL (OPT-AIL), which centers on performing online optimization for reward functions and optimism-regularized Bellman error minimization for Q-value functions. Theoretically, we prove that OPT-AIL achieves polynomial expert sample complexity and interaction complexity for learning near-expert policies. To our best knowledge, OPT-AIL is the first provably efficient AIL method with general function approximation. Practically, OPT-AIL only requires the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods in several challenging tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 被引用 2 次
- Hierarchical Value-Decomposed Offline Reinforcement Learning for Whole-Body ControlZhilong Zhang, Yunpeng Mei, Xinghao Du, Hongjie Cao 等ICLR 2026
- Improving Reward Model Generalization from Adversarial Process Enhanced PreferencesZhilong Zhang, Tian Xu, Xinghao Du, Xingchen Cao 等ICML 2025
- IL-SOAR : Imitation Learning with Soft Optimistic Actor cRiticStefano Viel, Luca Viano, Volkan CevherICML 2025
它引用的顶会 Paper26
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient AlgorithmsChi Jin, Qinghua Liu, Sobhan MiryoosefiNeurIPS 2021 · 被引用 264 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
相关 Paper
- Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence RateYufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran WangICML 2020 · 被引用 12 次
- Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation LearningTian Xu, Zexuan Chen, Zhilong Zhang, Yi-Chen Li 等ICML 2026
- Online Apprenticeship LearningLior Shani, Tom Zahavy, Shie MannorAAAI 2022 · 被引用 33 次
- On Computation and Generalization of Generative Adversarial Imitation LearningMinshuo Chen, Yizhou Wang, Tianyi Liu, Zhuoran Yang 等ICLR 2020 · 被引用 42 次
- Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow MechanismTian Xu, Chenyang Wang, Xiaochen Zhai, Ziniu Li 等ICML 2026
