Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate Policies
Ilyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao He
Abstract
Recently, the impressive empirical success of policy gradient (PG) methods has catalyzed the development of their theoretical foundations. Despite the huge efforts directed at the design of efficient stochastic PG-type algorithms, the understanding of their convergence to a globally optimal policy is still limited. In this work, we develop improved global convergence guarantees for a general class of Fisher-non-degenerate parameterized policies which allows to address the case of continuous state action spaces. First, we propose a Normalized Policy Gradient method with Implicit Gradient Transport (N-PG-IGT) and derive a sample complexity of this method for finding a global -optimal policy. Improving over the previously known complexity, this algorithm does not require the use of importance sampling or second-order information and samples only one trajectory per iteration. Second, we further improve this complexity to by considering a Hessian-Aided Recursive Policy Gradient ((N)-HARPG) algorithm enhanced with a correction based on a Hessian-vector product. Interestingly, both algorithms are simple and easy to implement: single-loop, do not require large batches of trajectories and sample at most two trajectories per iteration; computationally and memory efficient: they do not require expensive subroutines at each iteration and can be implemented with memory linear in the dimension of parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c1c00cc-65cf-47f2-b5bc-5baffa5b4904Cited by top-tier papers26
- Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive MethodsJunchi Yang, Xiang Li, Ilyas Fatkhullin, Niao HeNeurIPS 2023 · 33 citations
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 25 citations
- Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision ProcessesQinbo Bai, Washim Uddin Mondal, Vaneet AggarwalAAAI 2024 · 23 citations
- Reinforcement Learning with General Utilities: Simpler Variance Reduction and Large State-Action SpaceAnas Barakat, Ilyas Fatkhullin, Niao HeICML 2023 · 18 citations
- Sample-Efficient Constrained Reinforcement Learning with General ParameterizationWashim Uddin Mondal, Vaneet AggarwalNeurIPS 2024 · 15 citations
Related papers
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 128 citations
- Momentum-Based Policy Gradient MethodsFeihu Huang, Shangqian Gao, Jian Pei, Heng HuangICML 2020 · 47 citations
- Reusing Trajectories in Policy Gradients Enables Fast ConvergenceAlessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini et al.ICML 2026
- Learning Optimal Deterministic Policies with Stochastic Policy GradientsAlessandro Montenegro, Marco Mussi, Alberto Maria Metelli, Matteo PapiniICML 2024 · 11 citations
- Convergence Analysis of Policy Gradient Methods with Dynamic StochasticityAlessandro Montenegro, Marco Mussi, Matteo Papini, Alberto Maria MetelliICML 2025
