Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents
Woojun Kim, Yongjae Shin, Jongeui Park, Youngchul Sung
Abstract
Deep reinforcement learning (RL) has achieved remarkable success in solving complex tasks through its integration with deep neural networks (DNNs) as function approximators. However, the reliance on DNNs has introduced a new challenge called primacy bias, whereby these function approximators tend to prioritize early experiences, leading to overfitting. To mitigate this primacy bias, a reset method has been proposed, which performs periodic resets of a portion or the entirety of a deep RL agent while preserving the replay buffer. However, the use of the reset method can result in performance collapses after executing the reset, which can be detrimental from the perspective of safe RL and regret minimization. In this paper, we propose a new reset-based method that leverages deep ensemble learning to address the limitations of the vanilla reset method and enhance sample efficiency. The proposed method is evaluated through various experiments including those in the domain of safe RL. Numerical results show its effectiveness in high sample efficiency and safety considerations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f7dc55f-7b97-461e-a0f3-50f39d5a34feCited by top-tier papers7
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 5 citations
- Hard Tasks First: Multi-Task Reinforcement Learning Through Task SchedulingMyungsik Cho, Jongeui Park, Suyoung Lee, Youngchul SungICML 2024 · 4 citations
- Bayesian Ensemble for Sequential Decision-MakingRui Liu, Enmin Zhao, Lu Wang, Yu Li et al.ICLR 2026 · 2 citations
- ARS: Adaptive Reward Scaling for Multi-Task Reinforcement LearningMyungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul SungICML 2025
- Hyperspherical Normalization for Scalable Deep Reinforcement LearningHojoon Lee, Youngdo Lee, Takuma Seno, Donghu Kim et al.ICML 2025
Builds on7
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 168 citations
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 12 citations
Related papers
- Stay Hungry, Keep Learning: Sustainable Plasticity for Deep Reinforcement LearningHuaicheng Zhou, Zifeng Zhuang, Donglin WangICML 2025
- Efficient Deep Reinforcement Learning Requires Regulating OverfittingQiyang Li, Aviral Kumar, Ilya Kostrikov, Sergey LevineICLR 2023
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
- A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous ControlZilin Kang, Chenyuan Hu, Yu Luo, Zhecheng Yuan et al.ICML 2025
- Sample-Efficient Reinforcement Learning by Breaking the Replay Ratio BarrierPierluca D'Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon et al.ICLR 2023
