Maxmin Q-learning: Controlling the Estimation Bias of Q-learning
Qingfeng Lan, Yangchen Pan, Alona Fyshe, Martha White
摘要
Q-learning suffers from overestimation bias, because it approximates the maximum action value using the maximum estimated action value. Algorithms have been proposed to reduce overestimation bias, but we lack an understanding of how bias interacts with performance, and the extent to which existing algorithms mitigate bias. In this paper, we 1) highlight that the effect of overestimation bias on learning efficiency is environment-dependent; 2) propose a generalization of Q-learning, called Maxmin Q-learning, which provides a parameter to flexibly control bias; 3) show theoretically that there exists a parameter choice for Maxmin Q-learning that leads to unbiased estimation with a lower approximation variance than Q-learning; and 4) prove the convergence of our algorithm in the tabular case, as well as convergence of several previous Q-learning variants, using a novel Generalized Q-learning framework. We empirically verify that our algorithm better controls estimation bias in toy environments, and that it achieves superior performance on several benchmark problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 被引用 430 次
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi 等ICLR 2022 · 被引用 157 次
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 被引用 138 次
相关 Paper
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan 等ICML 2025
- Regularized Q-learning through Robust AveragingPeter Schmitt-Förster, Tobias SutterICML 2024
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 被引用 56 次
- Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action TasksHaobo Jiang, Jin Xie, Jian YangAAAI 2021 · 被引用 20 次
