Implementation Matters in Deep RL: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, Aleksander Madry
摘要
We study the roots of algorithmic progress in deep policy gradient algorithms through a case study on two popular algorithms, Proximal Policy Optimization and Trust Region Policy Optimization. We investigate the consequences of code-level optimizations: algorithm augmentations found only in implementations or described as auxiliary details to the core algorithm. Seemingly of secondary importance, such optimizations have a major impact on agent behavior. Our results show that they (a) are responsible for most of PPO's gain in cumulative reward over TRPO, and (b) fundamentally change how RL methods function. These insights show the difficulty, and importance, of attributing performance gains in deep reinforcement learning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper62
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye 等ICML 2024 · 被引用 274 次
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei 等ICLR 2026 · 被引用 168 次
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
相关 Paper
- Improving Value Estimation Critically Enhances Vanilla Policy GradientTao Wang, Ruipeng Zhang, Sicun GaoICML 2025
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning ImplementationsRajdeep Singh Hundal, Yan Xiao, Xiaochun Cao, Jin Song Dong 等ICSE 2025
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 41 次
- Differentiable Trust Region Layers for Deep Reinforcement LearningFabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Maria Ziesche 等ICLR 2021 · 被引用 23 次
- Simple Policy OptimizationZhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter 等ICML 2025
