Fractal Landscapes in Policy Optimization
Tao Wang, Sylvia L. Herbert, Sicun Gao
摘要
Policy gradient lies at the core of deep reinforcement learning (RL) in continuous domains. Despite much success, it is often observed in practice that RL training with policy gradient can fail for many reasons, even on standard control problems with known solutions. We propose a framework for understanding one inherent limitation of the policy gradient approach: the optimization landscape in the policy space can be extremely non-smooth or fractal for certain classes of MDPs, such that there does not exist gradient to be estimated in the first place. We draw on techniques from chaos theory and non-smooth analysis, and analyze the maximal Lyapunov exponents and Hölder exponents of the policy optimization objectives. Moreover, we develop a practical method that can estimate the local smoothness of objective function from samples to identify when the training process has encountered fractal landscapes. We show experiments to illustrate how some failure cases of policy optimization can be explained by such fractal landscapes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Enhancing Robustness in Deep Reinforcement Learning: A Lyapunov Exponent ApproachRory Young, Nicolas PugeaultNeurIPS 2024 · 被引用 5 次
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 被引用 2 次
- Improving Value Estimation Critically Enhances Vanilla Policy GradientTao Wang, Ruipeng Zhang, Sicun GaoICML 2025
它引用的顶会 Paper2
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 被引用 172 次
- Fractal Structure and Generalization Properties of Stochastic Optimization AlgorithmsAlexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban 等NeurIPS 2021 · 被引用 34 次
相关 Paper
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon 等NeurIPS 2023 · 被引用 11 次
- Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsShenao Zhang, Boyi Liu, Zhaoran Wang, Tuo ZhaoNeurIPS 2023 · 被引用 8 次
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li 等ICML 2023 · 被引用 19 次
- Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient EstimatorHaoyu Han, Heng YangICML 2026 · 被引用 3 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
