Stabilizing Dynamical Systems via Policy Gradient Methods
Juan C. Perdomo, Jack Umenberger, Max Simchowitz
摘要
Stabilizing an unknown control system is one of the most fundamental problems in control systems engineering. In this paper, we provide a simple, model-free algorithm for stabilizing fully observed dynamical systems. While model-free methods have become increasingly popular in practice due to their simplicity and flexibility, stabilization via direct policy search has received surprisingly little attention. Our algorithm proceeds by solving a series of discounted LQR problems, where the discount factor is gradually increased. We prove that this method efficiently recovers a stabilizing controller for linear systems, and for smooth, nonlinear systems within a neighborhood of their equilibria. Our approach overcomes a significant limitation of prior work, namely the need for a pre-given stabilizing control policy. We empirically evaluate the effectiveness of our approach on common control benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Global Convergence of Direct Policy Search for State-Feedback Robust Control: A Revisit of Nonsmooth Synthesis with Goldstein SubdifferentialXingang Guo, Bin HuNeurIPS 2022 · 被引用 16 次
- On the Sample Complexity of Stabilizing LTI Systems on a Single TrajectoryYang Hu, Adam Wierman, Guannan QuNeurIPS 2022 · 被引用 14 次
- Complexity of Derivative-Free Policy Optimization for Structured H∞ ControlXingang Guo, Darioush Keivan, Geir E. Dullerud, Peter J. Seiler 等NeurIPS 2023 · 被引用 12 次
- Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and AutoregressionAdam Block, Dylan J. Foster, Akshay Krishnamurthy, Max Simchowitz 等ICLR 2024 · 被引用 12 次
- Globally Optimal Policy Gradient Algorithms for Reinforcement Learning with PID Control PoliciesVipul Sharma, Wesley Suttle, S. SivaranjaniNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper2
相关 Paper
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with √T RegretAsaf B. Cassel, Tomer KorenICML 2021 · 被引用 20 次
- Globally Convergent Policy Search for Output EstimationJack Umenberger, Max Simchowitz, Juan C. Perdomo, Kaiqing Zhang 等NeurIPS 2022 · 被引用 16 次
- The Power of Learned Locally Linear Models for Nonlinear Policy OptimizationDaniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni 等ICML 2023 · 被引用 4 次
- Robust Reinforcement Learning: A Case Study in Linear Quadratic RegulationBo Pang, Zhong-Ping JiangAAAI 2021 · 被引用 43 次
- Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial StatesNoam Razin, Yotam Alexander, Edo Cohen-Karlik, Raja Giryes 等ICML 2024 · 被引用 5 次
