Globally Convergent Policy Search for Output Estimation
Jack Umenberger, Max Simchowitz, Juan C. Perdomo, Kaiqing Zhang, Russ Tedrake
摘要
We introduce the first direct policy search algorithm which provably converges to the globally optimal dynamic filter for the classical problem of predicting the outputs of a linear dynamical system, given noisy, partial observations. Despite the ubiquity of partial observability in practice, theoretical guarantees for direct policy search algorithms, one of the backbones of modern reinforcement learning, have proven difficult to achieve. This is primarily due to the degeneracies which arise when optimizing over filters that maintain an internal state. In this paper, we provide a new perspective on this challenging problem based on the notion of informativity, which intuitively requires that all components of a filter's internal state are representative of the true state of the underlying dynamical system. We show that informativity overcomes the aforementioned degeneracy. Specifically, we propose a regularizer which explicitly enforces informativity, and establish that gradient descent on this regularized objective -combined with a "reconditioning step" -converges to the globally optimal cost at a O(1/T ) rate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Data-driven Optimal Filtering for Linear Systems with Unknown Noise CovariancesShahriar Talebi, Amirhossein Taghvaei, Mehran MesbahiNeurIPS 2023 · 被引用 17 次
- Robot Fleet Learning via Policy MergingLirui Wang, Kaiqing Zhang, Allan Zhou, Max Simchowitz 等ICLR 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Stabilizing Dynamical Systems via Policy Gradient MethodsJuan C. Perdomo, Jack Umenberger, Max SimchowitzNeurIPS 2021 · 被引用 56 次
- Information is Power: Intrinsic Control via Information CaptureNicholas Rhinehart, Jenny Wang, Glen Berseth, John D. Co-Reyes 等NeurIPS 2021 · 被引用 14 次
- Scalable Online Exploration via CoverabilityPhilip Amortila, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 10 次
- Provably Efficient Reinforcement Learning in Partially Observable Dynamical SystemsMasatoshi Uehara, Ayush Sekhari, Jason D. Lee, Nathan Kallus 等NeurIPS 2022 · 被引用 48 次
- Global Convergence of Direct Policy Search for State-Feedback Robust Control: A Revisit of Nonsmooth Synthesis with Goldstein SubdifferentialXingang Guo, Bin HuNeurIPS 2022 · 被引用 16 次
