Lune

ICML2024Top-tier venue

Finite Time Logarithmic Regret Bounds for Self-Tuning Regulation

Rahul Singh, Akshay Mete, Avik Kar, Panganamala R. Kumar

2024Year

Abstract

We establish the first finite-time logarithmic regret bounds for the self-tuning regulation problem. We introduce a modified version of the certainty equivalence algorithm, which we call PIECE, that clips inputs in addition to utilizing probing inputs for exploration. We show that it has a C log T upper bound on the regret after T time-steps for bounded noise, and C log 3 T in the case of sub-Gaussian noise, unlike the LQ problem where logarithmic regret is shown to be not possible. The PIECE algorithm is also designed to address the critical challenge of poor initial transient performance of reinforcement learning algorithms for linear systems. Comparative simulation results illustrate the improved performance of PIECE.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 53b317a4-25a7-41cf-b8dc-1d62706e661e

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines