Generalized Proximal Policy Optimization with Sample Reuse
James Queeney, Yannis Paschalidis, Christos G. Cassandras
摘要
In real-world decision making tasks, it is critical for data-driven reinforcement learning methods to be both stable and sample efficient. On-policy methods typically generate reliable policy improvement throughout training, while off-policy methods make more efficient use of data through sample reuse. In this work, we combine the theoretically supported stability benefits of on-policy algorithms with the sample efficiency of off-policy algorithms. We develop policy improvement guarantees that are suitable for the off-policy setting, and connect these bounds to the clipping mechanism used in Proximal Policy Optimization. This motivates an off-policy version of the popular algorithm that we call Generalized Proximal Policy Optimization with Sample Reuse. We demonstrate both theoretically and empirically that our algorithm delivers improved performance by effectively balancing the competing goals of stability and sample efficiency. On-policy reinforcement learning methods such as Proximal Policy Optimization (PPO) [19] deliver stable performance throughout training due to their connection to theoretical policy improvement guarantees. These methods are motivated by a lower bound on the expected performance loss at every update, which can be approximated using samples generated by the current policy. The theoretically supported stability of these methods is very attractive, but the need for on-policy data and the highvariance nature of reinforcement learning often requires significant data to be collected between every update, resulting in high sample complexity and slow learning. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout ReplayYifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen 等NeurIPS 2025 · 被引用 63 次
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 41 次
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy TrainingYoussef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure 等ICLR 2026 · 被引用 39 次
- An Implicit Trust Region Approach to Behavior Regularized Offline Reinforcement LearningZhe Zhang, Xiaoyang TanAAAI 2024 · 被引用 10 次
- Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language ModelJing Liang, Jinyi Liu, Yi Ma, Hongyao Tang 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper2
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini 等ICLR 2021 · 被引用 52 次
相关 Paper
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 被引用 27 次
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- On Stationary Point Convergence of PPO-ClipRuinan Jin, Shuai Li, Baoxiang WangICLR 2024 · 被引用 14 次
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang 等ICML 2026 · 被引用 22 次
- Simple Policy OptimizationZhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter 等ICML 2025
