Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning Environments
Ryan Sullivan, J. K. Terry, Benjamin Black, John P. Dickerson
Abstract
Visualizing optimization landscapes has led to many fundamental insights in numeric optimization, and novel improvements to optimization techniques. However, visualizations of the objective that reinforcement learning optimizes (the "reward surface") have only ever been generated for a small number of narrow contexts. This work presents reward surfaces and related visualizations of 27 of the most widely used reinforcement learning environments in Gym for the first time. We also explore reward surfaces in the policy gradient direction and show for the first time that many popular reinforcement learning environments have frequent "cliffs" (sudden large drops in expected return). We demonstrate that A2C often "dives off" these cliffs into low reward regions of the parameter space while PPO avoids them, confirming a popular intuition for PPO's improved performance over previous methods. We additionally introduce a highly extensible library that allows researchers to easily generate these visualizations in the future. Our findings provide new intuition to explain the successes and failures of modern RL methods, and our visualizations concretely characterize several failure modes of reinforcement learning agents in novel ways.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon et al.NeurIPS 2023 · 11 citations
- Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Hao Wang, Devesh Tiwari, Jian Li et al.SC 2024 · 10 citations
- Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless ComputingHanfei Yu, Jian Li, Yang Hua, Xu Yuan et al.AAAI 2024 · 8 citations
- Identifying Policy Gradient SubspacesJan Schneider, Pierre Schumacher, Simon Guist, Le Chen et al.ICLR 2024 · 7 citations
- Nitro: Boosting Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Jacob Carter, Hao Wang, Devesh Tiwari et al.VLDB 2025 · 3 citations
Builds on4
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement LearningHiroki Furuta, Tatsuya Matsushima, Tadashi Kozuno, Yutaka Matsuo et al.ICML 2021 · 17 citations
Related papers
- POPGym: Benchmarking Partially Observable Reinforcement LearningSteven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki et al.ICLR 2023 · 5 citations
- Gray-Box Gaussian Processes for Automated Reinforcement LearningGresa Shala, André Biedenkapp, Frank Hutter, Josif GrabockaICLR 2023
- Flat Reward in Policy Parameter Space Implies Robust Reinforcement LearningHyun-Kyu Lee, Sung Whan YoonICLR 2025
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 27 citations
