Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
Mudit Gaur, Amrit S. Bedi, Di Wang, Vaneet Aggarwal
Abstract
The current state-of-the-art theoretical analysis of Actor-Critic (AC) algorithms significantly lags in addressing the practical aspects of AC implementations. This crucial gap needs bridging to bring the analysis in line with practical implementations of AC. To address this, we advocate for considering the MMCLG criteria: Multi-layer neural network parametrization for actor/critic, Markovian sampling, Continuous state-action spaces, the performance of the Last iterate, and Global optimality. These aspects are practically significant and have been largely overlooked in existing theoretical analyses of AC algorithms. In this work, we address these gaps by providing the first comprehensive theoretical analysis of AC algorithms that encompasses all five crucial practical aspects (covers MMCLG criteria). We establish global convergence sample complexity bounds of Õ ǫ -3 . We achieve this result through our novel use of the weak gradient domination property of MDP's and our unique analysis of the error in critic estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b4f52ff-94e1-41ce-934a-bb5db23208a0Cited by top-tier papers11
- On the Sample Complexity Bounds of Bilevel Reinforcement LearningMudit Gaur, Utsav Singh, Amrit Singh Bedi, Raghu Pasupathy et al.NeurIPS 2025 · 13 citations
- Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic AlgorithmYang Xu, Swetha Ganesh, Washim Uddin Mondal, Qinbo Bai et al.NeurIPS 2025 · 8 citations
- On the Convergence of Single-Timescale Actor-CriticNavdeep Kumar, Priyank Agrawal, Giorgia Ramponi, Kfir Y. Levy et al.NeurIPS 2025 · 4 citations
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo et al.ICML 2025
- Actor-Critics Can Achieve Optimal Sample EfficiencyKevin Tan, Wei Fan, Yuting WeiICML 2025
Builds on12
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 128 citations
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 110 citations
- Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate PoliciesIlyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao HeICML 2023 · 61 citations
Related papers
- Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network ApproximationXuyang Chen, Fengzhuo Zhang, Keyu Yan, Lin ZhaoICLR 2026
- Single-Timescale Actor-Critic Provably Finds Globally Optimal PolicyZuyue Fu, Zhuoran Yang, Zhaoran WangICLR 2021 · 52 citations
- Convergence of Actor-Critic with Multi-Layer Neural NetworksHaoxing Tian, Alex Olshevsky, Yannis PaschalidisNeurIPS 2023 · 13 citations
- Non-Asymptotic Analysis for Single-Loop (Natural) Actor-Critic with Compatible Function ApproximationYudan Wang, Yue Wang, Yi Zhou, Shaofeng ZouICML 2024 · 11 citations
- A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic ApproachSwetha Ganesh, Washim Uddin Mondal, Vaneet AggarwalICML 2025
