Regularized Q-Learning
Han-Dong Lim, Donghwan Lee
2024年份
1被引次数
2顶会引用
摘要
We consider a single-loop algorithm for regularized Q-learning with linear function approximation. The proposed algorithm is motivated by a bilevel optimization formulation of regularized Q-learning wherein the lower level optimization problem aims to identify a value function approximation that satisfies Bellman’s recursive optimality condition, and the upper level aims to find the projection onto the span of basis vectors. We show that under certain assumptions, the proposed algorithm converges to a stationary point in the presence of Markovian noise. In addition, we provide a performance guarantee for the policies derived from the proposed algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Adaptive Policy Backbone via Shared NetworkBumgeun Park, Donghwan LeeICML 2026
- Linear Q-Learning Does Not Diverge in L2: Convergence Rates to a Bounded SetXinyu Liu, Zixuan Xie, Shangtong ZhangICML 2025
它引用的顶会 Paper7
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta 等ICML 2020 · 被引用 49 次
- A Unified Switching System Perspective and Convergence Analysis of Q-Learning AlgorithmsDonghwan Lee, Niao HeNeurIPS 2020 · 被引用 48 次
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 被引用 39 次
相关 Paper
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 被引用 8 次
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 被引用 2 次
- Computationally Efficient RL under Linear Bellman Completeness for Deterministic DynamicsRunzhe Wu, Ayush Sekhari, Akshay Krishnamurthy, Wen SunICLR 2025
- Stabilizing Q-learning with Linear Architectures for Provable Efficient LearningAndrea Zanette, Martin J. WainwrightICML 2022 · 被引用 5 次
- Understanding and Leveraging Overparameterization in Recursive Value EstimationChenjun Xiao, Bo Dai, Jincheng Mei, Oscar A. Ramirez 等ICLR 2022 · 被引用 17 次
