On Stationary Point Convergence of PPO-Clip
Ruinan Jin, Shuai Li, Baoxiang Wang
摘要
Proximal policy optimization (PPO) has gained popularity in reinforcement learning (RL). Its PPO-Clip variant is one the most frequently implemented algorithms and is one of the first-to-try algorithms in RL tasks. This variant uses a clipped surrogate objective function not typically found in other algorithms. Many works have demonstrated the practical performance of PPO-Clip, but the theoretical understanding of it is limited to specific settings. In this work, we provide a comprehensive analysis that shows the stationary point convergence of PPO-Clip and the convergence rate thereof. Our analysis is new and overcomes many challenges, including the non-smooth nature of the clip operator, the potentially unbounded score function, and the involvement of the ratio of two stochastic policies. Our results and techniques might share new insights into PPO-Clip.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Tricks or Traps? A Deep Dive into RL for LLM ReasoningZihe Liu, Jiashun Liu, Yancheng He, Weixun Wang 等ICLR 2026 · 被引用 49 次
- On Entropy Control in LLM-RL AlgorithmsHan ShenICLR 2026 · 被引用 43 次
- From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy AssimilationZezhou Wang, Ziyun Zhang, Xiaoyi Zhang, Zhuzhong Qian 等ACL 2026 · 被引用 2 次
- Action-Dependent Optimality-Preserving Reward ShapingGrant C. Forbes, Jianxun Wang, Leonardo Villalobos-Arias, Arnav Jhala 等ICML 2025
它引用的顶会 Paper6
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári 等NeurIPS 2021 · 被引用 87 次
- Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate PoliciesIlyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao HeICML 2023 · 被引用 61 次
- Momentum-Based Policy Gradient MethodsFeihu Huang, Shangqian Gao, Jian Pei, Heng HuangICML 2020 · 被引用 47 次
相关 Paper
- PPO-Clip Attains Global Optimality: Towards Deeper Understandings of ClippingNai-Chieh Huang, Ping-Chun Hsieh, Kuo-Hao Ho, I-Chen WuAAAI 2024 · 被引用 34 次
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 被引用 27 次
- The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy MeasureXing Chen, Dongcui Diao, Hechang Chen, Hengshuai Yao 等AAAI 2023 · 被引用 28 次
- Generalized Proximal Policy Optimization with Sample ReuseJames Queeney, Yannis Paschalidis, Christos G. CassandrasNeurIPS 2021 · 被引用 80 次
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy TrainingYoussef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure 等ICLR 2026 · 被引用 39 次
