Dropout Q-Functions for Doubly Efficient Reinforcement Learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, Yoshimasa Tsuruoka
摘要
Randomized ensembled double Q-learning (REDQ) (Chen et al., 2021b) has recently achieved state-of-the-art sample efficiency on continuous-action reinforcement learning benchmarks. This superior sample efficiency is made possible by using a large Q-function ensemble. However, REDQ is much less computationally efficient than non-ensemble counterparts such as Soft Actor-Critic (SAC) (Haarnoja et al., 2018a). To make REDQ more computationally efficient, we propose a method of improving computational efficiency called DroQ, which is a variant of REDQ that uses a small ensemble of dropout Q-functions. Our dropout Q-functions are simple Q-functions equipped with dropout connection and layer normalization. Despite its simplicity of implementation, our experimental results indicate that DroQ is doubly (sample and computationally) efficient. It achieved comparable sample efficiency with REDQ, much better computational efficiency than REDQ, and comparable computational efficiency with that of SAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 被引用 153 次
- Revisiting the Minimalist Approach to Offline Reinforcement LearningDenis Tarasov, Vladislav Kurenkov, Alexander Nikulin, Sergey KolesnikovNeurIPS 2023 · 被引用 148 次
- Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous controlMichal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos 等NeurIPS 2024 · 被引用 119 次
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus 等ICLR 2024 · 被引用 106 次
它引用的顶会 Paper11
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
相关 Paper
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 被引用 26 次
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner 等ICLR 2026 · 被引用 20 次
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
- Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep NetworksLitian Liang, Yaosheng Xu, Stephen McAleer, Dailin Hu 等ICML 2022 · 被引用 25 次
- Towards Inference Efficient Deep Ensemble LearningZiyue Li, Kan Ren, Yifan Yang, Xinyang Jiang 等AAAI 2023 · 被引用 18 次
