Learning Optimal Deterministic Policies with Stochastic Policy Gradients
Alessandro Montenegro, Marco Mussi, Alberto Maria Metelli, Matteo Papini
Abstract
Policy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters. Stochastic controllers, however, are often undesirable from a practical perspective because of their lack of robustness, safety, and traceability. In common practice, stochastic (hyper)policies are learned only to deploy their deterministic version. In this paper, we make a step towards the theoretical understanding of this practice. After introducing a novel framework for modeling this scenario, we study the global convergence to the best deterministic policy, under (weak) gradient domination assumptions. Then, we illustrate how to tune the exploration level used for learning to optimize the trade-off between the sample complexity and the performance of the deployed deterministic policy. Finally, we quantitatively compare action-based and parameter-based exploration, giving a formal guise to intuitive results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6473e65-716d-4ced-a876-a234b9ed2323Cited by top-tier papers4
- Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement LearningAlessandro Montenegro, Marco Mussi, Matteo Papini, Alberto Maria MetelliNeurIPS 2024 · 3 citations
- Globally Optimal Policy Gradient Algorithms for Reinforcement Learning with PID Control PoliciesVipul Sharma, Wesley Suttle, S. SivaranjaniNeurIPS 2025 · 3 citations
- Convergence Analysis of Policy Gradient Methods with Dynamic StochasticityAlessandro Montenegro, Marco Mussi, Matteo Papini, Alberto Maria MetelliICML 2025
- Reusing Trajectories in Policy Gradients Enables Fast ConvergenceAlessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini et al.ICML 2026
Builds on8
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 128 citations
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate PoliciesIlyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao HeICML 2023 · 61 citations
Related papers
- On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement LearningAnas Barakat, Souradip Chakraborty, Peihong Yu, Pratap Tokekar et al.NeurIPS 2025 · 6 citations
- Reparameterized Policy Learning for Multimodal Trajectory OptimizationZhiao Huang, Litian Liang, Zhan Ling, Xuanlin Li et al.ICML 2023 · 21 citations
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans et al.NeurIPS 2025 · 2 citations
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsSeohong Park, Jaekyeom Kim, Gunhee KimNeurIPS 2021 · 33 citations
