Sample Complexity of Policy Gradient Finding Second-Order Stationary Points
Long Yang, Qian Zheng, Gang Pan
Abstract
The policy-based reinforcement learning (RL) can be considered as maximization of its objective. However, due to the inherent non-concavity of its objective, the policy gradient method to a first-order stationary point (FOSP) cannot guar- antee a maximal point. A FOSP can be a minimal or even a saddle point, which is undesirable for RL. It has be found that if all the saddle points are strict, all the second-order station- ary points (SOSP) are exactly equivalent to local maxima. Instead of FOSP, we consider SOSP as the convergence criteria to characterize the sample complexity of policy gradient. Our result shows that policy gradient converges to an (ε, √εχ)-SOSP with probability at least 1 − O(δ) after the total cost of O(ε−9/2)sinificantly improves the state of the art cost O(ε−9).Our analysis is based on the key idea that decomposes the parameter space Rp into three non-intersected regions: non-stationary point region, saddle point region, and local optimal region, then making a local improvement of the objective of RL in each region. This technique can be potentially generalized to extensive policy gradient methods. For the complete proof, please refer to https://arxiv.org/pdf/2012.01491.pdf.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8f8bed0-d9b7-484e-ab9b-88410b71d845Cited by top-tier papers7
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang et al.NeurIPS 2022 · 95 citations
- Policy Optimization with Stochastic Mirror DescentLong Yang, Yu Zhang, Gang Zheng, Qian Zheng et al.AAAI 2022 · 38 citations
- Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time AnalysisZiyi Chen, Yi Zhou, Rong-Rong Chen, Shaofeng ZouICML 2022 · 35 citations
- Controlling Type Confounding in Ad Hoc Teamwork with Instance-wise Teammate Feedback RectificationDong Xing, Pengjie Gu, Qian Zheng, Xinrun Wang et al.ICML 2023 · 4 citations
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 4 citations
Builds on1
Related papers
- On the convergence of policy gradient methods to Nash equilibria in general stochastic gamesAngeliki Giannou, Kyriakos Lotidis, Panayotis Mertikopoulos, Emmanouil V. Vlatakis-GkaragkounisNeurIPS 2022 · 28 citations
- Variational Policy Gradient Method for Reinforcement Learning with General UtilitiesJunyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvári et al.NeurIPS 2020 · 170 citations
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Nearly Optimal Policy Optimization with Stable at Any Time GuaranteeTianhao Wu, Yunchang Yang, Han Zhong, Liwei Wang et al.ICML 2022 · 15 citations
- Low-Switching Policy Gradient with Exploration via Online Sensitivity SamplingYunfan Li, Yiran Wang, Yu Cheng, Lin YangICML 2023 · 6 citations
