Lune

NeurIPS2024Top-tier venue

Regularized Q-Learning

Han-Dong Lim, Donghwan Lee

2024Year
1Citations
2Top-tier citations

Abstract

We consider a single-loop algorithm for regularized Q-learning with linear function approximation. The proposed algorithm is motivated by a bilevel optimization formulation of regularized Q-learning wherein the lower level optimization problem aims to identify a value function approximation that satisfies Bellman’s recursive optimality condition, and the upper level aims to find the projection onto the span of basis vectors. We show that under certain assumptions, the proposed algorithm converges to a stationary point in the presence of Markovian noise. In addition, we provide a performance guarantee for the policies derived from the proposed algorithm.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7d388a8b-b342-45fd-ba1e-3f4a106d660e

Cited by top-tier papers2

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines