The Role of Baselines in Policy Gradient Optimization
Jincheng Mei, Wesley Chung, Valentin Thomas, Bo Dai, Csaba Szepesvári, Dale Schuurmans
摘要
We study the effect of baselines in on-policy stochastic policy gradient optimization, and close the gap between the theory and practice of policy optimization methods. Our first contribution is to show that the state value baseline allows on-policy stochastic natural policy gradient (NPG) to converge to a globally optimal policy at an rate, which was not previously known. The analysis relies on two novel findings: the expected progress of the NPG update satisfies a stochastic version of the non-uniform ojasiewicz (N) inequality, and with probability 1 the state value baseline prevents the optimal action's probability from vanishing, thus ensuring sufficient exploration. Importantly, these results provide a new understanding of the role of baselines in stochastic policy gradient: by showing that the variance of natural policy gradient estimates remains unbounded with or without a baseline, we find that variance reduction cannot explain their utility in this setting. Instead, the analysis reveals that the primary effect of the value baseline is to reduce the aggressiveness of the updates rather than their variance. That is, we demonstrate that a finite variance is not necessary for almost sure convergence of stochastic NPG, while controlling update aggressiveness is both necessary and sufficient. Additional experimental results verify these theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataFahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov 等ICML 2024 · 被引用 189 次
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin 等ICML 2024 · 被引用 165 次
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM ReasoningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalICLR 2026 · 被引用 7 次
- Stochastic Gradient Succeeds for BanditsJincheng Mei, Zixin Zhong, Bo Dai, Alekh Agarwal 等ICML 2023 · 被引用 6 次
- Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning ratesJincheng Mei, Bo Dai, Alekh Agarwal, Sharan Vaswani 等NeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper8
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári 等NeurIPS 2021 · 被引用 87 次
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
- Leveraging Non-uniformity in First-order Non-convex OptimizationJincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvári 等ICML 2021 · 被引用 55 次
相关 Paper
- Beyond Variance Reduction: Understanding the True Impact of Baselines on Policy OptimizationWesley Chung, Valentin Thomas, Marlos C. Machado, Nicolas Le RouxICML 2021 · 被引用 35 次
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 被引用 128 次
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans 等NeurIPS 2025 · 被引用 2 次
- Understanding the Effect of Stochasticity in Policy OptimizationJincheng Mei, Bo Dai, Chenjun Xiao, Csaba Szepesvári 等NeurIPS 2021 · 被引用 24 次
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 被引用 26 次
