Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
Nate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon, Marc G. Bellemare
Abstract
Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective on these behaviors by studying the return landscape: the mapping between a policy and a return. We find that popular algorithms traverse noisy neighborhoods of this landscape, in which a single update to the policy parameters leads to a wide range of returns. By taking a distributional view of these returns, we map the landscape, characterizing failure-prone regions of policy space and revealing a hidden dimension of policy quality. We show that the landscape exhibits surprising structure by finding simple paths in parameter space which improve the stability of a policy. To conclude, we develop a distribution-aware procedure which finds such paths, navigating away from noisy neighborhoods in order to improve the robustness of a policy. Taken together, our results provide new insight into the optimization, evaluation, and design of agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95939c9f-6575-4cc0-b802-07566de627d4Cited by top-tier papers6
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu et al.ICLR 2024 · 97 citations
- Do Transformer World Models Give Better Policy Gradients?Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro et al.ICML 2024 · 7 citations
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman et al.ICLR 2026 · 6 citations
- Enhancing Robustness in Deep Reinforcement Learning: A Lyapunov Exponent ApproachRory Young, Nicolas PugeaultNeurIPS 2024 · 5 citations
- SHAPO: Sharpness-Aware Policy Optimization for Safe ExplorationKaustubh Mani, Yann Pequignot, Vincent Mai, Liam PaullICLR 2026 · 3 citations
Builds on6
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny et al.ICLR 2020 · 99 citations
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 38 citations
Related papers
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 10 citations
- Flat Reward in Policy Parameter Space Implies Robust Reinforcement LearningHyun-Kyu Lee, Sung Whan YoonICLR 2025
- Understanding and Diagnosing Deep Reinforcement LearningEzgi KorkmazICML 2024 · 10 citations
- Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region ApproachJames Queeney, Ioannis Ch. Paschalidis, Christos G. CassandrasAAAI 2021 · 11 citations
- Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust DecisionsEzgi Korkmaz, Jonah Brown-CohenICML 2023 · 16 citations
