Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error Feedback
Hang Wang, Sen Lin, Junshan Zhang
Abstract
The ensemble method is a promising way to mitigate the overestimation issue in Q-learning, where multiple function approximators are used to estimate the action values. It is known that the estimation bias hinges heavily on the ensemble size (i.e., the number of Q-function approximators used in the target), and that determining the 'right' ensemble size is highly nontrivial, because of the time-varying nature of the function approximation errors during the learning process. To tackle this challenge, we first derive an upper bound and a lower bound on the estimation bias, based on which the ensemble size is adapted to drive the bias to be nearly zero, thereby coping with the impact of the time-varying approximation errors accordingly. Motivated by the theoretic findings, we advocate that the ensemble method can be combined with Model Identification Adaptive Control (MIAC) for effective ensemble size adaptation. Specifically, we devise Adaptive Ensemble Q-learning (AdaEQ), a generalized ensemble method with two key steps: (a) approximation error characterization which serves as the feedback for flexibly controlling the ensemble size, and (b) ensemble size adaptation tailored towards minimizing the estimation bias. Extensive experiments are carried out to show that AdaEQ can improve the learning performance than the existing methods for the MuJoCo benchmark. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e130b73-9687-4747-982f-b6043b1f195cCited by top-tier papers6
- Stochastic Q-learning for Large Discrete Action SpacesFares Fourati, Vaneet Aggarwal, Mohamed-Slim AlouiniICML 2024 · 9 citations
- Double Gumbel Q-LearningDavid Yu-Tung Hui, Aaron C. Courville, Pierre-Luc BaconNeurIPS 2023 · 8 citations
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision ProcessesDavid Ireland, Giovanni MontanaICLR 2024 · 6 citations
- Multi-timescale Reinforcement Learning by Value ReconstructionZhan Su, Peixi Peng, Xinyu Hu, Cong Li et al.ICML 2026
Builds on6
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 26 citations
- On the Power of Randomization for Scheduling Real-Time Traffic in Wireless NetworksChristos Tsanikidis, Javad GhaderiINFOCOM 2020 · 24 citations
Related papers
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan et al.ICML 2025
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- The Surprising Difficulty of Search in Model-Based Reinforcement LearningWei-Di Chang, Mikael Henaff, Brandon Amos, Gregory Dudek et al.ICML 2026 · 4 citations
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 4 citations
