The Power of Learned Locally Linear Models for Nonlinear Policy Optimization
Daniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni, Stephen Tu
摘要
A common pipeline in learning-based control is to iteratively estimate a model of system dynamics, and apply a trajectory optimization algorithm - e.g. - on the learned model to minimize a target cost. This paper conducts a rigorous analysis of a simplified variant of this strategy for general nonlinear systems. We analyze an algorithm which iterates between estimating local linear models of nonlinear system dynamics and performing -like policy updates. We demonstrate that this algorithm attains sample complexity polynomial in relevant problem parameters, and, by synthesizing locally stabilizing gains, overcomes exponential dependence in problem horizon. Experimental results validate the performance of our algorithm, and compare to natural deep-learning baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Provable Guarantees for Generative Behavior Cloning: Bridging Low-Level Stability and High-Level BehaviorAdam Block, Ali Jadbabaie, Daniel Pfrommer, Max Simchowitz 等NeurIPS 2023 · 被引用 44 次
- Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and AutoregressionAdam Block, Dylan J. Foster, Akshay Krishnamurthy, Max Simchowitz 等ICLR 2024 · 被引用 12 次
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with √T RegretAsaf B. Cassel, Tomer KorenICML 2021 · 被引用 20 次
- Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPsSiow Meng Low, Akshat Kumar, Scott SannerAAAI 2022 · 被引用 3 次
- Enforcing robust control guarantees within neural network policiesPriya L. Donti, Melrose Roderick, Mahyar Fazlyab, J. Zico KolterICLR 2021 · 被引用 12 次
- Stabilizing Dynamical Systems via Policy Gradient MethodsJuan C. Perdomo, Jack Umenberger, Max SimchowitzNeurIPS 2021 · 被引用 56 次
- Optimal Exploration for Model-Based RL in Nonlinear SystemsAndrew Wagenmaker, Guanya Shi, Kevin JamiesonNeurIPS 2023 · 被引用 29 次
