Stabilizing Dynamical Systems via Policy Gradient Methods
Juan C. Perdomo, Jack Umenberger, Max Simchowitz
Abstract
Stabilizing an unknown control system is one of the most fundamental problems in control systems engineering. In this paper, we provide a simple, model-free algorithm for stabilizing fully observed dynamical systems. While model-free methods have become increasingly popular in practice due to their simplicity and flexibility, stabilization via direct policy search has received surprisingly little attention. Our algorithm proceeds by solving a series of discounted LQR problems, where the discount factor is gradually increased. We prove that this method efficiently recovers a stabilizing controller for linear systems, and for smooth, nonlinear systems within a neighborhood of their equilibria. Our approach overcomes a significant limitation of prior work, namely the need for a pre-given stabilizing control policy. We empirically evaluate the effectiveness of our approach on common control benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Global Convergence of Direct Policy Search for State-Feedback Robust Control: A Revisit of Nonsmooth Synthesis with Goldstein SubdifferentialXingang Guo, Bin HuNeurIPS 2022 · 16 citations
- On the Sample Complexity of Stabilizing LTI Systems on a Single TrajectoryYang Hu, Adam Wierman, Guannan QuNeurIPS 2022 · 14 citations
- Complexity of Derivative-Free Policy Optimization for Structured H∞ ControlXingang Guo, Darioush Keivan, Geir E. Dullerud, Peter J. Seiler et al.NeurIPS 2023 · 12 citations
- Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and AutoregressionAdam Block, Dylan J. Foster, Akshay Krishnamurthy, Max Simchowitz et al.ICLR 2024 · 12 citations
- Globally Optimal Policy Gradient Algorithms for Reinforcement Learning with PID Control PoliciesVipul Sharma, Wesley Suttle, S. SivaranjaniNeurIPS 2025 · 3 citations
Builds on2
Related papers
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with √T RegretAsaf B. Cassel, Tomer KorenICML 2021 · 20 citations
- Globally Convergent Policy Search for Output EstimationJack Umenberger, Max Simchowitz, Juan C. Perdomo, Kaiqing Zhang et al.NeurIPS 2022 · 16 citations
- The Power of Learned Locally Linear Models for Nonlinear Policy OptimizationDaniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni et al.ICML 2023 · 4 citations
- Robust Reinforcement Learning: A Case Study in Linear Quadratic RegulationBo Pang, Zhong-Ping JiangAAAI 2021 · 43 citations
- Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial StatesNoam Razin, Yotam Alexander, Edo Cohen-Karlik, Raja Giryes et al.ICML 2024 · 5 citations
