Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
Sara Klein, Simon Weissmann, Leif Döring
Abstract
Markov Decision Processes (MDPs) are a formal framework for modeling and solving sequential decision-making problems. In finite-time horizons such problems are relevant for instance for optimal stopping or specific supply chain problems, but also in the training of large language models. In contrast to infinite horizon MDPs optimal policies are not stationary, policies must be learned for every single epoch. In practice all parameters are often trained simultaneously, ignoring the inherent structure suggested by dynamic programming. This paper introduces a combination of dynamic programming and policy gradient called dynamic policy gradient, where the parameters are trained backwards in time. For the tabular softmax parametrisation we carry out the convergence analysis for simultaneous and dynamic policy gradient towards global optima, both in the exact and sampled gradient settings without regularisation. It turns out that the use of dynamic policy gradient training much better exploits the structure of finite- time problems which is reflected in improved convergence bounds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- On Entropy Control in LLM-RL AlgorithmsHan ShenICLR 2026 · 43 citations
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy ShapingShuang Chen, Hangyu Guo, Yimeng Ye, Shijue Huang et al.ICLR 2026 · 23 citations
- Does Stochastic Gradient really succeed for bandits?Dorian Baudry, Emmeran Johnson, Simon Vary, Ciara Pike-Burke et al.NeurIPS 2025 · 3 citations
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans et al.NeurIPS 2025 · 2 citations
- Structure Matters: Dynamic Policy GradientSara Klein, Xiangyuan Zhang, Tamer Basar, Simon Weissmann et al.NeurIPS 2025 · 1 citation
Builds on10
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate PoliciesIlyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao HeICML 2023 · 61 citations
- Leveraging Non-uniformity in First-order Non-convex OptimizationJincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvári et al.ICML 2021 · 55 citations
- Momentum-Based Policy Gradient MethodsFeihu Huang, Shangqian Gao, Jian Pei, Heng HuangICML 2020 · 47 citations
Related papers
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regimeAndrea Agazzi, Jianfeng LuICLR 2021 · 3 citations
- Global Convergence of Policy Gradient in Average Reward MDPsNavdeep Kumar, Yashaswini Murthy, Itai Shufaro, Kfir Yehuda Levy et al.ICLR 2025
- Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled RegularizationSafwan Labbi, Daniil Tiapkin, Paul Mangold, Eric MoulinesICLR 2026 · 2 citations
- Doubly Regularized Markov Decision Processes for Robust Reinforcement LearningYiting He, Zhishuai Liu, Pan XuICML 2026 · 213 citations
