Ensemble Bootstrapping for Q-Learning
Oren Peer, Chen Tessler, Nadav Merlis, Ron Meir
摘要
Q-learning (QL), a common reinforcement learning algorithm, suffers from over-estimation bias due to the maximization term in the optimal Bellman operator. This bias may lead to sub-optimal behavior. Double-Q-learning tackles this issue by utilizing two estimators, yet results in an under-estimation bias. Similar to over-estimation in Q-learning, in certain scenarios, the under-estimation bias may degrade performance. In this work, we introduce a new bias-reduced algorithm called Ensemble Bootstrapped Q-Learning (EBQL), a natural extension of Double-Q-learning to ensembles. We analyze our method both theoretically and empirically. Theoretically, we prove that EBQL-like updates yield lower MSE when estimating the maximal mean of a set of independent random variables. Empirically, we show that there exist domains where both over and under-estimation result in sub-optimal performance. Finally, We demonstrate the superior performance of a deep RL variant of EBQL over other deep QL algorithms for a suite of ATARI games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Ensemble-based Deep Reinforcement Learning for Vehicle Routing Problems under Distribution ShiftYuan Jiang, Zhiguang Cao, Yaoxin Wu, Wen Song 等NeurIPS 2023 · 被引用 43 次
- Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error FeedbackHang Wang, Sen Lin, Junshan ZhangNeurIPS 2021 · 被引用 27 次
- Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep NetworksLitian Liang, Yaosheng Xu, Stephen McAleer, Dailin Hu 等ICML 2022 · 被引用 25 次
- Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticTianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan 等ICML 2024 · 被引用 23 次
- The Curse of Diversity in Ensemble-Based ExplorationZhixuan Lin, Pierluca D'Oro, Evgenii Nikishin, Aaron C. CourvilleICLR 2024 · 被引用 9 次
它引用的顶会 Paper2
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
相关 Paper
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- On the Estimation Bias in Double Q-LearningZhizhou Ren, Guangxiang Zhu, Hao Hu, Beining Han 等NeurIPS 2021 · 被引用 35 次
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan 等ICML 2025
- Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action TasksHaobo Jiang, Jin Xie, Jian YangAAAI 2021 · 被引用 20 次
- Regularized Q-learning through Robust AveragingPeter Schmitt-Förster, Tobias SutterICML 2024
