ADDQ: Adaptive distributional double Q-learning
Leif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan, Martin Slowik
Abstract
Bias problems in the estimation of maxima of random variables are a well-known obstacle that drastically slows down Q-learning algorithms. We propose to use additional insight gained from distributional reinforcement learning to deal with the overestimation in a locally adaptive way. This helps to combine the strengths and weaknesses of the different Q-learning variants in a unified framework. Our framework ADDQ is simple to implement, existing RL algorithms can be improved with a few lines of additional code. We provide experimental results in tabular, Atari, and MuJoCo environments for discrete and continuous control problems, comparisons with state-of-the-art methods, and a proof of convergence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e9c43ab-e922-46f2-836a-894b4ae3b8e9Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 26 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
Related papers
- Regularized Q-learning through Robust AveragingPeter Schmitt-Förster, Tobias SutterICML 2024
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 22 citations
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 4 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li et al.AAAI 2022 · 20 citations
