Finite Time Logarithmic Regret Bounds for Self-Tuning Regulation
Rahul Singh, Akshay Mete, Avik Kar, Panganamala R. Kumar
摘要
We establish the first finite-time logarithmic regret bounds for the self-tuning regulation problem. We introduce a modified version of the certainty equivalence algorithm, which we call PIECE, that clips inputs in addition to utilizing probing inputs for exploration. We show that it has a C log T upper bound on the regret after T time-steps for bounded noise, and C log 3 T in the case of sub-Gaussian noise, unlike the LQ problem where logarithmic regret is shown to be not possible. The PIECE algorithm is also designed to address the critical challenge of poor initial transient performance of reinforcement learning algorithms for linear systems. Comparative simulation results illustrate the improved performance of PIECE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 被引用 209 次
- Logarithmic Regret Bound in Partially Observable Linear Dynamical SystemsSahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima AnandkumarNeurIPS 2020 · 被引用 106 次
- Logarithmic Regret for Learning Linear Quadratic Regulators EfficientlyAsaf B. Cassel, Alon Cohen, Tomer KorenICML 2020 · 被引用 68 次
- Augmented RBMLE-UCB Approach for Adaptive Control of Linear Quadratic SystemsAkshay Mete, Rahul Singh, P. R. KumarNeurIPS 2022 · 被引用 10 次
相关 Paper
- Task-Optimal Exploration in Linear Dynamical SystemsAndrew J. Wagenmaker, Max Simchowitz, Kevin JamiesonICML 2021 · 被引用 24 次
- Regret Bounds for Episodic Risk-Sensitive Linear Quadratic RegulatorWenhao Xu, Xuefeng Gao, Xuedong HeICLR 2025
- Efficient Optimistic Exploration in Linear-Quadratic Regulators via Lagrangian RelaxationMarc Abeille, Alessandro LazaricICML 2020 · 被引用 31 次
- Improved Worst-Case Regret Bounds for Randomized Least-Squares Value IterationPriyank Agrawal, Jinglin Chen, Nan JiangAAAI 2021 · 被引用 24 次
- Near-Optimal Randomized Exploration for Tabular Markov Decision ProcessesZhihan Xiong, Ruoqi Shen, Qiwen Cui, Maryam Fazel 等NeurIPS 2022 · 被引用 17 次
