Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems
Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima Anandkumar
摘要
We study the problem of system identification and adaptive control in partially observable linear dynamical systems. Adaptive and closed-loop system identification is a challenging problem due to correlations introduced in data collection. In this paper, we present the first model estimation method with finite-time guarantees in both open and closed-loop system identification. Deploying this estimation method, we propose adaptive control online learning (AdaptOn), an efficient reinforcement learning algorithm that adaptively learns the system dynamics and continuously updates its controller through online learning steps. AdaptOn estimates the model dynamics by occasionally solving a linear regression problem through interactions with the environment. Using policy re-parameterization and the estimated model, AdaptOn constructs counterfactual loss functions to be used for updating the controller through online gradient descent. Over time, AdaptOn improves its model estimates and obtains more accurate gradient updates to improve the controller. We show that AdaptOn achieves a regret upper bound of , after time steps of agent-environment interaction. To the best of our knowledge, AdaptOn is the first algorithm that achieves regret in adaptive control of unknown partially observable linear dynamical systems which includes linear quadratic Gaussian (LQG) control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Online Optimization with Memory and Competitive ControlGuanya Shi, Yiheng Lin, Soon-Jo Chung, Yisong Yue 等NeurIPS 2020 · 被引用 66 次
- Meta-Adaptive Nonlinear Control: Theory and AlgorithmsGuanya Shi, Kamyar Azizzadenesheli, Michael O'Connell, Soon-Jo Chung 等NeurIPS 2021 · 被引用 62 次
- Making Non-Stochastic Control (Almost) as Easy as StochasticMax SimchowitzNeurIPS 2020 · 被引用 44 次
- Online Control of Unknown Time-Varying Dynamical SystemsEdgar Minasyan, Paula Gradu, Max Simchowitz, Elad HazanNeurIPS 2021 · 被引用 38 次
- Learning the Linear Quadratic Regulator from Nonlinear ObservationsZakaria Mhammedi, Dylan J. Foster, Max Simchowitz, Dipendra Misra 等NeurIPS 2020 · 被引用 33 次
它引用的顶会 Paper3
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 被引用 209 次
- Logarithmic Regret for Adversarial Online ControlDylan J. Foster, Max SimchowitzICML 2020 · 被引用 82 次
- Logarithmic Regret for Learning Linear Quadratic Regulators EfficientlyAsaf B. Cassel, Alon Cohen, Tomer KorenICML 2020 · 被引用 68 次
相关 Paper
- Finite Sample Analyses for Continuous-time Linear Systems: System Identification and Online ControlHongyi Zhou, Jingwei Li, Jingzhao ZhangNeurIPS 2025
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with √T RegretAsaf B. Cassel, Tomer KorenICML 2021 · 被引用 20 次
- Sample Efficient Reinforcement Learning with Partial Dynamics KnowledgeMeshal Alharbi, Mardavij Roozbehani, Munther A. DahlehAAAI 2024 · 被引用 4 次
- Regret Bounds for Episodic Risk-Sensitive Linear Quadratic RegulatorWenhao Xu, Xuefeng Gao, Xuedong HeICLR 2025
- Provably Efficient Model-based Policy AdaptationYuda Song, Aditi Mavalankar, Wen Sun, Sicun GaoICML 2020 · 被引用 11 次
