Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
Yu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang, Huazhe Xu, Xianyuan Zhan
Abstract
Off-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning. However, most existing off-policy RL algorithms fail to maximally exploit the information in the replay buffer, limiting sample efficiency and policy performance. In this work, we discover that concurrently training an offline RL policy based on the shared online replay buffer can sometimes outperform the original online learning policy, though the occurrence of such performance gains remains uncertain. This motivates a new possibility of harnessing the emergent outperforming offline optimal policy to improve online policy learning. Based on this insight, we present Offline-Boosted Actor-Critic (OBAC), a model-free online RL framework that elegantly identifies the outperforming offline policy through value comparison, and uses it as an adaptive constraint to guarantee stronger policy learning performance. Our experiments demonstrate that OBAC outperforms other popular model-free RL baselines and rivals advanced model-based RL methods in terms of sample efficiency and asymptotic performance across 53 tasks spanning 6 task suites 1 . Introduction Online model-free deep reinforcement learning (RL) methods have achieved success in many challenging sequential
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eff461f1-9b39-4f4d-bd2a-faacb00e1a10Cited by top-tier papers5
- Flow-Based Policy for Online Reinforcement LearningLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun et al.NeurIPS 2025 · 39 citations
- Value Improved Actor Critic AlgorithmsYaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok et al.NeurIPS 2025 · 7 citations
- A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous ControlZilin Kang, Chenyuan Hu, Yu Luo, Zhecheng Yuan et al.ICML 2025
- Debiased Model-based Representations for Sample-efficient Continuous ControlJiafei Lyu, Zichuan Lin, Scott Fujimoto, Kai Yang et al.ICML 2026
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
Builds on32
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
Related papers
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song et al.AAAI 2023 · 47 citations
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 1 citation
- Model-Based Offline Meta-Reinforcement Learning with RegularizationSen Lin, Jialin Wan, Tengyu Xu, Yingbin Liang et al.ICLR 2022 · 20 citations
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
