Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Haoyu Han, Heng Yang
摘要
Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the noise-to-signal ratio (NSR) of a policy-gradient estimator, defined as the estimator variance (noise) normalized by the squared norm of the true gradient (signal). Our main result is that, for (i) finite-horizon linear systems with Gaussian policies and linear state-feedback, and (ii) finite-horizon polynomial systems with Gaussian policies and polynomial feedback, the NSR of the REINFORCE estimator can be characterized exactly—either in closed form or via numerical moment-evaluation algorithms—without approximation. For general nonlinear dynamics and expressive policies (including neural policies), we further derive a general upper bound on the variance. These characterizations enable a direct examination of how NSR varies across policy parameters and how it evolves along optimization trajectories (e.g. SGD and Adam). Across a range of examples, we find that the NSR landscape is highly non-uniform and typically increases as the policy approaches an optimum; in some regimes it blows up, which can trigger training instability and policy collapse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 被引用 153 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
- Beyond Variance Reduction: Understanding the True Impact of Baselines on Policy OptimizationWesley Chung, Valentin Thomas, Marlos C. Machado, Nicolas Le RouxICML 2021 · 被引用 35 次
相关 Paper
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 被引用 10 次
- Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsShenao Zhang, Boyi Liu, Zhaoran Wang, Tuo ZhaoNeurIPS 2023 · 被引用 8 次
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 被引用 2 次
- On the convergence of policy gradient methods to Nash equilibria in general stochastic gamesAngeliki Giannou, Kyriakos Lotidis, Panayotis Mertikopoulos, Emmanouil V. Vlatakis-GkaragkounisNeurIPS 2022 · 被引用 28 次
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian SamplingHuaqing Xiong, Tengyu Xu, Yingbin Liang, Wei ZhangAAAI 2021 · 被引用 37 次
