Phasic Policy Gradient
Karl Cobbe, Jacob Hilton, Oleg Klimov, John Schulman
Abstract
We introduce Phasic Policy Gradient (PPG), a reinforcement learning framework which modifies traditional on-policy actor-critic methods by separating policy and value function training into distinct phases. In prior methods, one must choose between using a shared network or separate networks to represent the policy and value function. Using separate networks avoids interference between objectives, while using a shared network allows useful features to be shared. PPG is able to achieve the best of both worlds by splitting optimization into two phases, one that advances training and one that distills features. PPG also enables the value function to be more aggressively optimized with a higher level of sample reuse. Compared to PPO, we find that PPG significantly improves sample efficiency on the challenging Procgen Benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1deaeeb-90df-41a3-b611-b9b6935bc339Cited by top-tier papers55
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Learning to drive from a world on railsDian Chen, Vladlen Koltun, Philipp KrähenbühlICCV 2021 · 164 citations
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 116 citations
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 72 citations
Builds on1
Related papers
- PPG Reloaded: An Empirical Study on What Matters in Phasic Policy GradientKaixin Wang, Daquan Zhou, Jiashi Feng, Shie MannorICML 2023 · 1 citation
- Rethinking Value Function Learning for Generalization in Reinforcement LearningSeungyong Moon, JunYeong Lee, Hyun Oh SongNeurIPS 2022 · 17 citations
- Studying the Interplay Between the Actor and Critic Representations in Reinforcement LearningSamuel Garcin, Trevor McInroe, Pablo Samuel Castro, Christopher G. Lucas et al.ICLR 2025
- SAPG: Split and Aggregate Policy GradientsJayesh Singla, Ananye Agarwal, Deepak PathakICML 2024 · 19 citations
- DNA: Proximal Policy Optimization with a Dual Network ArchitectureMatthew Aitchison, Penny SweetserNeurIPS 2022 · 7 citations
