Ordering-based Conditions for Global Convergence of Policy Gradient Methods
Jincheng Mei, Bo Dai, Alekh Agarwal, Mohammad Ghavamzadeh, Csaba Szepesvári, Dale Schuurmans
摘要
We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the policy update and the representation. textcolorblueFirst, we establish a few key observations that frame the study: (i) Global convergence can be achieved under linear function approximation without policy or reward realizability, both for the standard Softmax PG and natural policy gradient (NPG). (ii) Approximation error is not a key quantity for characterizing global convergence in either algorithm. (iii) The conditions on the representation that imply global convergence are different between these two algorithms. Overall, these observations call into question approximation error as an appropriate quantity for characterizing the global convergence of PG methods under linear function approximation. blueSecond, motivated by these observations, we establish new general results: (i) NPG with linear function approximation achieves global convergence if and only if the projection of the reward onto the representable space preserves the optimal action's rank, a quantity that is not strongly related to approximation error. (ii) The global convergence of Softmax PG occurs if the representation satisfies a non-domination condition and can preserve the ranking of rewards, which goes well beyond policy or reward realizability. We provide experimental results to support these theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient LearningAlekh Agarwal, Mikael Henaff, Sham M. Kakade, Wen SunNeurIPS 2020 · 被引用 126 次
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
- The Role of Baselines in Policy Gradient OptimizationJincheng Mei, Wesley Chung, Valentin Thomas, Bo Dai 等NeurIPS 2022 · 被引用 34 次
- Linear Convergence of Natural Policy Gradient Methods with Log-Linear PoliciesRui Yuan, Simon Shaolei Du, Robert M. Gower, Alessandro Lazaric 等ICLR 2023 · 被引用 1 次
相关 Paper
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 被引用 128 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- ϕ-Update: A Class of Policy Update Methods with Policy Convergence GuaranteeWenye Li, Jiacai Liu, Ke WeiICLR 2025
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans 等NeurIPS 2025 · 被引用 2 次
- Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning ratesJincheng Mei, Bo Dai, Alekh Agarwal, Sharan Vaswani 等NeurIPS 2024 · 被引用 5 次
