Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
Nate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon, Marc G. Bellemare
摘要
Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective on these behaviors by studying the return landscape: the mapping between a policy and a return. We find that popular algorithms traverse noisy neighborhoods of this landscape, in which a single update to the policy parameters leads to a wide range of returns. By taking a distributional view of these returns, we map the landscape, characterizing failure-prone regions of policy space and revealing a hidden dimension of policy quality. We show that the landscape exhibits surprising structure by finding simple paths in parameter space which improve the stability of a policy. To conclude, we develop a distribution-aware procedure which finds such paths, navigating away from noisy neighborhoods in order to improve the robustness of a policy. Taken together, our results provide new insight into the optimization, evaluation, and design of agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- Do Transformer World Models Give Better Policy Gradients?Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro 等ICML 2024 · 被引用 7 次
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman 等ICLR 2026 · 被引用 6 次
- Enhancing Robustness in Deep Reinforcement Learning: A Lyapunov Exponent ApproachRory Young, Nicolas PugeaultNeurIPS 2024 · 被引用 5 次
- SHAPO: Sharpness-Aware Policy Optimization for Safe ExplorationKaustubh Mani, Yann Pequignot, Vincent Mai, Liam PaullICLR 2026 · 被引用 3 次
它引用的顶会 Paper6
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny 等ICLR 2020 · 被引用 99 次
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 被引用 38 次
相关 Paper
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 被引用 10 次
- Flat Reward in Policy Parameter Space Implies Robust Reinforcement LearningHyun-Kyu Lee, Sung Whan YoonICLR 2025
- Understanding and Diagnosing Deep Reinforcement LearningEzgi KorkmazICML 2024 · 被引用 10 次
- Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region ApproachJames Queeney, Ioannis Ch. Paschalidis, Christos G. CassandrasAAAI 2021 · 被引用 11 次
- Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust DecisionsEzgi Korkmaz, Jonah Brown-CohenICML 2023 · 被引用 16 次
