Regularized Q-learning through Robust Averaging
Peter Schmitt-Förster, Tobias Sutter
Abstract
We propose a new Q-learning variant, called 2RA Q-learning, that addresses some weaknesses of existing Q-learning methods in a principled manner. One such weakness is an underlying estimation bias which cannot be controlled and often results in poor performance. We propose a distributionally robust estimator for the maximum expected value term, which allows us to precisely control the level of estimation bias introduced. The distributionally robust estimator admits a closed-form solution such that the proposed algorithm has a computational cost per iteration comparable to Watkins' Q-learning. For the tabular case, we show that 2RA Q-learning converges to the optimal policy and analyze its asymptotic mean-squared error. Lastly, we conduct numerical experiments for various settings, which corroborate our theoretical findings and indicate that 2RA Q-learning often performs better than existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 417d59d1-691f-490a-819a-00c87a9f5f12Builds on6
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong et al.ICML 2022 · 72 citations
- A Unified Switching System Perspective and Convergence Analysis of Q-Learning AlgorithmsDonghwan Lee, Niao HeNeurIPS 2020 · 48 citations
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 26 citations
Related papers
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan et al.ICML 2025
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 22 citations
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 4 citations
- Single-Trajectory Distributionally Robust Reinforcement LearningZhipeng Liang, Xiaoteng Ma, José H. Blanchet, Jun Yang et al.ICML 2024 · 15 citations
