Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents
Woojun Kim, Yongjae Shin, Jongeui Park, Youngchul Sung
摘要
Deep reinforcement learning (RL) has achieved remarkable success in solving complex tasks through its integration with deep neural networks (DNNs) as function approximators. However, the reliance on DNNs has introduced a new challenge called primacy bias, whereby these function approximators tend to prioritize early experiences, leading to overfitting. To mitigate this primacy bias, a reset method has been proposed, which performs periodic resets of a portion or the entirety of a deep RL agent while preserving the replay buffer. However, the use of the reset method can result in performance collapses after executing the reset, which can be detrimental from the perspective of safe RL and regret minimization. In this paper, we propose a new reset-based method that leverages deep ensemble learning to address the limitations of the vanilla reset method and enhance sample efficiency. The proposed method is evaluated through various experiments including those in the domain of safe RL. Numerical results show its effectiveness in high sample efficiency and safety considerations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 被引用 5 次
- Hard Tasks First: Multi-Task Reinforcement Learning Through Task SchedulingMyungsik Cho, Jongeui Park, Suyoung Lee, Youngchul SungICML 2024 · 被引用 4 次
- Bayesian Ensemble for Sequential Decision-MakingRui Liu, Enmin Zhao, Lu Wang, Yu Li 等ICLR 2026 · 被引用 2 次
- ARS: Adaptive Reward Scaling for Multi-Task Reinforcement LearningMyungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul SungICML 2025
- Hyperspherical Normalization for Scalable Deep Reinforcement LearningHojoon Lee, Youngdo Lee, Takuma Seno, Donghu Kim 等ICML 2025
它引用的顶会 Paper7
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 被引用 168 次
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 被引用 12 次
相关 Paper
- Stay Hungry, Keep Learning: Sustainable Plasticity for Deep Reinforcement LearningHuaicheng Zhou, Zifeng Zhuang, Donglin WangICML 2025
- Efficient Deep Reinforcement Learning Requires Regulating OverfittingQiyang Li, Aviral Kumar, Ilya Kostrikov, Sergey LevineICLR 2023
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 被引用 56 次
- A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous ControlZilin Kang, Chenyuan Hu, Yu Luo, Zhecheng Yuan 等ICML 2025
- Sample-Efficient Reinforcement Learning by Breaking the Replay Ratio BarrierPierluca D'Oro, Max Schwarzer, Evgenii Nikishin, Pierre-Luc Bacon 等ICLR 2023
