Augmented RBMLE-UCB Approach for Adaptive Control of Linear Quadratic Systems
Akshay Mete, Rahul Singh, P. R. Kumar
Abstract
We consider the problem of controlling an unknown stochastic linear system with quadratic costs - called the adaptive LQ control problem. We re-examine an approach called ''Reward Biased Maximum Likelihood Estimate'' (RBMLE) that was proposed more than forty years ago, and which predates the ''Upper Confidence Bound'' (UCB) method as well as the definition of ''regret'' for bandit problems. It simply added a term favoring parameters with larger rewards to the criterion for parameter estimation. We show how the RBMLE and UCB methods can be reconciled, and thereby propose an Augmented RBMLE-UCB algorithm that combines the penalty of the RBMLE method with the constraints of the UCB method, uniting the two approaches to optimism in the face of uncertainty. We establish that theoretically, this method retains regret, the best-known so far. We further compare the empirical performance of the proposed Augmented RBMLE-UCB and the standard RBMLE (without the augmentation) with UCB, Thompson Sampling, Input Perturbation, Randomized Certainty Equivalence and StabL on many real-world examples including flight control of Boeing 747 and Unmanned Aerial Vehicle. We perform extensive simulation studies showing that the Augmented RBMLE consistently outperforms UCB, Thompson Sampling and StabL by a huge margin, while it is marginally better than Input Perturbation and moderately better than Randomized Certainty Equivalence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cf57e42-8b37-4732-82fb-976dd4212bbfCited by top-tier papers5
- Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationZhihan Liu, Miao Lu, Wei Xiong, Han Zhong et al.NeurIPS 2023 · 30 citations
- Learning the Uncertainty Sets of Linear Control Systems via Set Membership: A Non-asymptotic AnalysisYingying Li, Jing Yu, Lauren E. Conger, Taylan Kargin et al.ICML 2024 · 13 citations
- Bayesian Optimistic Optimization: Optimistic Exploration for Model-based Reinforcement LearningChenyang Wu, Tianci Li, Zongzhang Zhang, Yang YuNeurIPS 2022 · 9 citations
- Robust System Identification: Finite-sample Guarantees and Connection to RegularizationHyuk Park, Grani A. Hanasusanto, Yingying LiICLR 2025
- Finite Time Logarithmic Regret Bounds for Self-Tuning RegulationRahul Singh, Akshay Mete, Avik Kar, Panganamala R. KumarICML 2024
Builds on3
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 209 citations
- Reward-Biased Maximum Likelihood Estimation for Linear Stochastic BanditsYu-Heng Hung, Ping-Chun Hsieh, Xi Liu, P. R. KumarAAAI 2021 · 16 citations
- Exploration Through Reward Biasing: Reward-Biased Maximum Likelihood Estimation for Stochastic Multi-Armed BanditsXi Liu, Ping-Chun Hsieh, Yu-Heng Hung, Anirban Bhattacharya et al.ICML 2020 · 16 citations
Related papers
- UCB Momentum Q-learning: Correcting the bias without forgettingPierre Ménard, Omar Darwiche Domingues, Xuedong Shang, Michal ValkoICML 2021 · 53 citations
- Reward-Biased Maximum Likelihood Estimation for Neural Contextual Bandits: A Distributional Learning PerspectiveYu-Heng Hung, Ping-Chun HsiehAAAI 2023 · 2 citations
- Robust Contextual Bandits via BootstrappingQiao Tang, Hong Xie, Yunni Xia, Jia Lee et al.AAAI 2021 · 2 citations
- Budgeted Multi-Armed Bandits with Asymmetric Confidence IntervalsMarco Heyden, Vadim Arzamasov, Edouard Fouché, Klemens BöhmKDD 2024 · 1 citation
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 66 citations
