The Power of Learned Locally Linear Models for Nonlinear Policy Optimization
Daniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni, Stephen Tu
Abstract
A common pipeline in learning-based control is to iteratively estimate a model of system dynamics, and apply a trajectory optimization algorithm - e.g. - on the learned model to minimize a target cost. This paper conducts a rigorous analysis of a simplified variant of this strategy for general nonlinear systems. We analyze an algorithm which iterates between estimating local linear models of nonlinear system dynamics and performing -like policy updates. We demonstrate that this algorithm attains sample complexity polynomial in relevant problem parameters, and, by synthesizing locally stabilizing gains, overcomes exponential dependence in problem horizon. Experimental results validate the performance of our algorithm, and compare to natural deep-learning baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73144acd-4432-448d-a7f2-2b0dcd9d954cCited by top-tier papers3
- Provable Guarantees for Generative Behavior Cloning: Bridging Low-Level Stability and High-Level BehaviorAdam Block, Ali Jadbabaie, Daniel Pfrommer, Max Simchowitz et al.NeurIPS 2023 · 44 citations
- Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and AutoregressionAdam Block, Dylan J. Foster, Akshay Krishnamurthy, Max Simchowitz et al.ICLR 2024 · 12 citations
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 2 citations
Builds on1
Related papers
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with √T RegretAsaf B. Cassel, Tomer KorenICML 2021 · 20 citations
- Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPsSiow Meng Low, Akshat Kumar, Scott SannerAAAI 2022 · 3 citations
- Enforcing robust control guarantees within neural network policiesPriya L. Donti, Melrose Roderick, Mahyar Fazlyab, J. Zico KolterICLR 2021 · 12 citations
- Stabilizing Dynamical Systems via Policy Gradient MethodsJuan C. Perdomo, Jack Umenberger, Max SimchowitzNeurIPS 2021 · 56 citations
- Optimal Exploration for Model-Based RL in Nonlinear SystemsAndrew Wagenmaker, Guanya Shi, Kevin JamiesonNeurIPS 2023 · 29 citations
