Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
Safwan Labbi, Daniil Tiapkin, Paul Mangold, Eric Moulines
Abstract
Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditioned optimization landscapes and lead to exponentially slow convergence. Although this can be mitigated by preconditioning, this solution is often computationally expensive. Instead, we propose replacing the softmax with an alternative family of policy parameterizations based on the generalized -. We further advocate coupling this parameterization with a regularizer induced by the same -divergence, which improves the optimization landscape and ensures that the resulting regularized objective satisfies a Polyak--Łojasiewicz inequality. Leveraging this structure, we establish the for stochastic policy gradient methods for finite MDPs . We also derive sample-complexity bounds for the unregularized problem and show that -PG, with Tsallis divergences achieves in contrast to the exponential complexity incurred by the standard softmax parameterization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Emergence of Exploration in Policy Gradient Reinforcement Learning via RetryingSoichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno et al.ICML 2026 · 6 citations
- Refined Analysis of Entropy-Regularized Actor-CriticSafwan Labbi, Paul Mangold, Daniil Tiapkin, Eric MoulinesICML 2026
Builds on5
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 162 citations
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li et al.NeurIPS 2020 · 56 citations
- Loss Functions and Operators Generated by f-DivergencesVincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander et al.ICML 2025
Related papers
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 23 citations
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 2 citations
- On the Global Convergence Rates of Decentralized Softmax Gradient Play in Markov Potential GamesRunyu Zhang, Jincheng Mei, Bo Dai, Dale Schuurmans et al.NeurIPS 2022 · 38 citations
- f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-DivergencesSiddhant Agarwal, Ishan Durugkar, Peter Stone, Amy ZhangNeurIPS 2023 · 22 citations
- Best of Both Worlds Policy OptimizationChristoph Dann, Chen-Yu Wei, Julian ZimmertICML 2023 · 17 citations
