Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Haoyu Han, Heng Yang
Abstract
Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the noise-to-signal ratio (NSR) of a policy-gradient estimator, defined as the estimator variance (noise) normalized by the squared norm of the true gradient (signal). Our main result is that, for (i) finite-horizon linear systems with Gaussian policies and linear state-feedback, and (ii) finite-horizon polynomial systems with Gaussian policies and polynomial feedback, the NSR of the REINFORCE estimator can be characterized exactly—either in closed form or via numerical moment-evaluation algorithms—without approximation. For general nonlinear dynamics and expressive policies (including neural policies), we further derive a general upper bound on the variance. These characterizations enable a direct examination of how NSR varies across policy parameters and how it evolves along optimization trajectories (e.g. SGD and Adam). Across a range of examples, we find that the NSR landscape is highly non-uniform and typically increases as the policy approaches an optimum; in some regimes it blows up, which can trigger training instability and policy collapse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cec3de2-b657-4f58-8023-b68e0670c651Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 153 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Beyond Variance Reduction: Understanding the True Impact of Baselines on Policy OptimizationWesley Chung, Valentin Thomas, Marlos C. Machado, Nicolas Le RouxICML 2021 · 35 citations
Related papers
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 10 citations
- Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsShenao Zhang, Boyi Liu, Zhaoran Wang, Tuo ZhaoNeurIPS 2023 · 8 citations
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 2 citations
- On the convergence of policy gradient methods to Nash equilibria in general stochastic gamesAngeliki Giannou, Kyriakos Lotidis, Panayotis Mertikopoulos, Emmanouil V. Vlatakis-GkaragkounisNeurIPS 2022 · 28 citations
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian SamplingHuaqing Xiong, Tengyu Xu, Yingbin Liang, Wei ZhangAAAI 2021 · 37 citations
