A Closer Look at Deep Policy Gradients
Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, Aleksander Madry
Abstract
We study how the behavior of deep policy gradient algorithms reflects the conceptual framework motivating their development. To this end, we propose a fine-grained analysis of state-of-the-art methods based on key elements of this framework: gradient estimation, value prediction, and optimization landscapes. Our results show that the behavior of deep policy gradient algorithms often deviates from what their motivating framework would predict: the surrogate objective does not match the true reward landscape, learned value estimators fail to fit the true value function, and gradient estimates poorly correlate with the "true" gradient. The mismatch between predicted and empirical behavior we uncover highlights our poor understanding of current methods, and indicates the need to move beyond current benchmark-centric evaluation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b076a4d4-48e9-4dca-8576-df193276f84aCited by top-tier papers24
- Recomposing the Reinforcement Learning Building Blocks with HypernetworksElad Sarafian, Shai Keynan, Sarit KrausICML 2021 · 42 citations
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy OptimizationPierluca D'Oro, Wojciech JaskowskiNeurIPS 2020 · 33 citations
- On Proximal Policy Optimization's Heavy-tailed GradientsSaurabh Garg, Joshua Zhanson, Emilio Parisotto, Adarsh Prasad et al.ICML 2021 · 32 citations
- The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy MeasureXing Chen, Dongcui Diao, Hechang Chen, Hengshuai Yao et al.AAAI 2023 · 28 citations
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon et al.NeurIPS 2025 · 26 citations
Related papers
- Improving Deep Policy Gradients with Value Function SearchEnrico Marchesini, Christopher AmatoICLR 2023
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 10 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Identifying Policy Gradient SubspacesJan Schneider, Pierre Schumacher, Simon Guist, Le Chen et al.ICLR 2024 · 7 citations
