On the Estimation Bias in Double Q-Learning
Zhizhou Ren, Guangxiang Zhu, Hao Hu, Beining Han, Jianglun Chen, Chongjie Zhang
Abstract
Double Q-learning is a classical method for reducing overestimation bias, which is caused by taking maximum estimated values in the Bellman operation. Its variants in the deep Q-learning paradigm have shown great promise in producing reliable value prediction and improving learning performance. However, as shown by prior work, double Q-learning is not fully unbiased and suffers from underestimation bias. In this paper, we show that such underestimation bias may lead to multiple non-optimal fixed points under an approximate Bellman operator. To address the concerns of converging to non-optimal stationary solutions, we propose a simple but effective approach as a partial fix for the underestimation bias in double Q-learning. This approach leverages an approximate dynamic programming to bound the target value. We extensively evaluate our proposed method in the Atari benchmark tasks and demonstrate its significant improvement over baseline algorithms. an interesting fact that, under the effects of approximation error, double Q-learning may have multiple non-optimal fixed points. The main cause of such non-optimal fixed points is the underestimation bias of double Q-learning. Regarding this issue, we provide some analysis to characterize what kind of Bellman operators may suffer from the same problem, and how the agent may behave around these fixed points. To address the potential risk of converging to non-optimal solutions, we propose doubly bounded Q-learning to reduce the underestimation in double Q-learning. The main idea of this approach is to leverage an abstracted dynamic programming as a second value estimator to rule out non-optimal fixed points. The experiments show that the proposed method has shown great promise in improving
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Episodic Reinforcement Learning with Associative MemoryGuangxiang Zhu, Zichuan Lin, Guangwen Yang, Chongjie ZhangICLR 2020 · 56 citations
- Generalizable Episodic Memory for Deep Reinforcement LearningHao Hu, Jianing Ye, Guangxiang Zhu, Zhizhou Ren et al.ICML 2021 · 44 citations
- On the Expressivity of Neural Networks for Deep Reinforcement LearningKefan Dong, Yuping Luo, Tianhe Yu, Chelsea Finn et al.ICML 2020 · 33 citations
Related papers
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 22 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action TasksHaobo Jiang, Jin Xie, Jian YangAAAI 2021 · 20 citations
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li et al.AAAI 2022 · 20 citations
- Finite-Time Analysis for Double Q-learningHuaqing Xiong, Lin Zhao, Yingbin Liang, Wei ZhangNeurIPS 2020 · 33 citations
