Regularized Q-Learning
Han-Dong Lim, Donghwan Lee
Abstract
We consider a single-loop algorithm for regularized Q-learning with linear function approximation. The proposed algorithm is motivated by a bilevel optimization formulation of regularized Q-learning wherein the lower level optimization problem aims to identify a value function approximation that satisfies Bellman’s recursive optimality condition, and the upper level aims to find the projection onto the span of basis vectors. We show that under certain assumptions, the proposed algorithm converges to a stationary point in the presence of Markovian noise. In addition, we provide a performance guarantee for the policies derived from the proposed algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d388a8b-b342-45fd-ba1e-3f4a106d660eCited by top-tier papers2
- Adaptive Policy Backbone via Shared NetworkBumgeun Park, Donghwan LeeICML 2026
- Linear Q-Learning Does Not Diverge in L2: Convergence Rates to a Bounded SetXinyu Liu, Zixuan Xie, Shangtong ZhangICML 2025
Builds on7
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta et al.ICML 2020 · 49 citations
- A Unified Switching System Perspective and Convergence Analysis of Q-Learning AlgorithmsDonghwan Lee, Niao HeNeurIPS 2020 · 48 citations
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 39 citations
Related papers
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 8 citations
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 2 citations
- Computationally Efficient RL under Linear Bellman Completeness for Deterministic DynamicsRunzhe Wu, Ayush Sekhari, Akshay Krishnamurthy, Wen SunICLR 2025
- Stabilizing Q-learning with Linear Architectures for Provable Efficient LearningAndrea Zanette, Martin J. WainwrightICML 2022 · 5 citations
- Understanding and Leveraging Overparameterization in Recursive Value EstimationChenjun Xiao, Bo Dai, Jincheng Mei, Oscar A. Ramirez et al.ICLR 2022 · 17 citations
