Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
Safwan Labbi, Daniil Tiapkin, Paul Mangold, Eric Moulines
摘要
Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditioned optimization landscapes and lead to exponentially slow convergence. Although this can be mitigated by preconditioning, this solution is often computationally expensive. Instead, we propose replacing the softmax with an alternative family of policy parameterizations based on the generalized -. We further advocate coupling this parameterization with a regularizer induced by the same -divergence, which improves the optimization landscape and ensures that the resulting regularized objective satisfies a Polyak--Łojasiewicz inequality. Leveraging this structure, we establish the for stochastic policy gradient methods for finite MDPs . We also derive sample-complexity bounds for the unregularized problem and show that -PG, with Tsallis divergences achieves in contrast to the exponential complexity incurred by the standard softmax parameterization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Emergence of Exploration in Policy Gradient Reinforcement Learning via RetryingSoichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno 等ICML 2026 · 被引用 6 次
- Refined Analysis of Entropy-Regularized Actor-CriticSafwan Labbi, Paul Mangold, Daniil Tiapkin, Eric MoulinesICML 2026
它引用的顶会 Paper5
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
- Loss Functions and Operators Generated by f-DivergencesVincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander 等ICML 2025
相关 Paper
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 被引用 23 次
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 被引用 2 次
- On the Global Convergence Rates of Decentralized Softmax Gradient Play in Markov Potential GamesRunyu Zhang, Jincheng Mei, Bo Dai, Dale Schuurmans 等NeurIPS 2022 · 被引用 38 次
- f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-DivergencesSiddhant Agarwal, Ishan Durugkar, Peter Stone, Amy ZhangNeurIPS 2023 · 被引用 22 次
- Best of Both Worlds Policy OptimizationChristoph Dann, Chen-Yu Wei, Julian ZimmertICML 2023 · 被引用 17 次
