Lune

AAAI2021Top-tier venue

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

Priyank Agrawal, Jinglin Chen, Nan Jiang

2021Year
24Citations
12Top-tier citations

Abstract

This paper studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our Õ(H 2 S √ AT ) high-probability worst-case regret bound improves the previous sharpest worst-case regret bounds for RLSVI and matches the existing state-of-the-art worst-case TS-based regret bounds. * These two authors contributed equally. 1 Õ (•) hides dependence on logarithmic factors.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d964fffb-f76f-4818-a2d2-4237f2e55ec2

Cited by top-tier papers12

Ask how each one uses it

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines