Lune

NeurIPS2025Top-tier venue

REINFORCE Converges to Optimal Policies with Any Learning Rate

Samuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans, Csaba Szepesvári, Jincheng Mei

2025Year
2Citations

Abstract

We prove that the classic REINFORCE stochastic policy gradient (SPG) method converges to globally optimal policies in finite-horizon Markov Decision Processes (MDPs) with any constant learning rate. To avoid the need for small or decaying learning rates, we introduce two key innovations in the stochastic bandit setting, which we then extend to MDPs. First , we identify a new exploration property: the online SPG method samples every action infinitely often, improving on previous results that only guaranteed at least two actions would be sampled infinitely often. This means SPG inherently achieves asymptotic exploration without modification. Second , we eliminate the assumption of unique mean reward values, a condition that previous convergence analyses in the bandit setting relied on, but that does not translate to MDPs. Our results deepen the theoretical understanding of SPG in both bandit problems and MDPs, with a focus on how it handles the exploration-exploitation trade-off when standard analysis techniques for optimization and stochastic approximation methods cannot be applied, as is the case with large constant learning rates.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d92e7d1e-0fff-467e-8eae-eae6874c9d3c

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines