Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
Tianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan, Jianwei Zhang, Huazhe Xu
Abstract
Learning high-quality -value functions plays a key role in the success of many modern off-policy deep reinforcement learning (RL) algorithms. Previous works primarily focus on addressing the value overestimation issue, an outcome of adopting function approximators and off-policy learning. Deviating from the common viewpoint, we observe that -values are often underestimated in the latter stage of the RL training process, potentially hindering policy learning and reducing sample efficiency. We find that such a long-neglected phenomenon is often related to the use of inferior actions from the current policy in Bellman updates as compared to the more optimal action samples in the replay buffer. To address this issue, our insight is to incorporate sufficient exploitation of past successes while maintaining exploration optimism. We propose the Blended Exploitation and Exploration (BEE) operator, a simple yet effective approach that updates -value using both historical best-performing actions and the current policy. Based on BEE, the resulting practical algorithm BAC outperforms state-of-the-art methods in over 50 continuous control tasks and achieves strong performance in failure-prone scenarios and real-world robot tasks. Benchmark results and videos are available at https://jity16.github.io/BEE/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caf46bfd-7250-4e9b-b4b4-d17ff7daaa4dCited by top-tier papers12
- DrM: Mastering Visual Reinforcement Learning through Dormant Ratio MinimizationGuowei Xu, Ruijie Zheng, Yongyuan Liang, Xiyao Wang et al.ICLR 2024 · 53 citations
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 48 citations
- Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement LearningMichal Nauman, Michal Bortkiewicz, Piotr Milos, Tomasz Trzcinski et al.ICML 2024 · 46 citations
- Flow-Based Policy for Online Reinforcement LearningLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun et al.NeurIPS 2025 · 39 citations
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo et al.ICML 2024 · 20 citations
Builds on20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
Related papers
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel et al.NeurIPS 2021 · 75 citations
- Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningMotoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta et al.ICML 2025
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 69 citations
- A Perspective of Q-value Estimation on Offline-to-Online Reinforcement LearningYinmin Zhang, Jie Liu, Chuming Li, Yazhe Niu et al.AAAI 2024 · 28 citations
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 38 citations
