Phasic Policy Gradient
Karl Cobbe, Jacob Hilton, Oleg Klimov, John Schulman
摘要
We introduce Phasic Policy Gradient (PPG), a reinforcement learning framework which modifies traditional on-policy actor-critic methods by separating policy and value function training into distinct phases. In prior methods, one must choose between using a shared network or separate networks to represent the policy and value function. Using separate networks avoids interference between objectives, while using a shared network allows useful features to be shared. PPG is able to achieve the best of both worlds by splitting optimization into two phases, one that advances training and one that distills features. PPG also enables the value function to be more aggressively optimized with a higher level of sample reuse. Compared to PPO, we find that PPG significantly improves sample efficiency on the challenging Procgen Benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Learning to drive from a world on railsDian Chen, Vladlen Koltun, Philipp KrähenbühlICCV 2021 · 被引用 164 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 被引用 72 次
它引用的顶会 Paper1
相关 Paper
- PPG Reloaded: An Empirical Study on What Matters in Phasic Policy GradientKaixin Wang, Daquan Zhou, Jiashi Feng, Shie MannorICML 2023 · 被引用 1 次
- Rethinking Value Function Learning for Generalization in Reinforcement LearningSeungyong Moon, JunYeong Lee, Hyun Oh SongNeurIPS 2022 · 被引用 17 次
- Studying the Interplay Between the Actor and Critic Representations in Reinforcement LearningSamuel Garcin, Trevor McInroe, Pablo Samuel Castro, Christopher G. Lucas 等ICLR 2025
- SAPG: Split and Aggregate Policy GradientsJayesh Singla, Ananye Agarwal, Deepak PathakICML 2024 · 被引用 19 次
- DNA: Proximal Policy Optimization with a Dual Network ArchitectureMatthew Aitchison, Penny SweetserNeurIPS 2022 · 被引用 7 次
