Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement Learning
Shenzhi Wang, Qisen Yang, Jiawei Gao, Matthieu Gaetan Lin, Hao Chen, Liwei Wu, Ning Jia, Shiji Song, Gao Huang
Abstract
Offline-to-online reinforcement learning (RL) is a training paradigm that combines pre-training on a pre-collected dataset with fine-tuning in an online environment. However, the incorporation of online fine-tuning can intensify the well-known distributional shift problem. Existing solutions tackle this problem by imposing a policy constraint on the policy improvement objective in both offline and online learning. They typically advocate a single balance between policy improvement and constraints across diverse data collections. This one-size-fits-all manner may not optimally leverage each collected sample due to the significant variation in data quality across different states. To this end, we introduce Family Offline-to-Online RL (FamO2O), a simple yet effective framework that empowers existing algorithms to determine state-adaptive improvement-constraint balances. FamO2O utilizes a universal model to train a family of policies with different improvement/constraint intensities, and a balance model to select a suitable policy for each state. Theoretically, we prove that state-adaptive balances are necessary for achieving a higher policy performance upper bound. Empirically, extensive experiments show that FamO2O offers a statistically significant improvement over various existing methods, achieving state-of-the-art performance on the D4RL benchmark. Codes are available at https://github.com/LeapLabTHU/FamO2O .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f3d947e-99c1-4d12-8bbf-df1b7e790552Cited by top-tier papers20
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 114 citations
- A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware PerspectiveYunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen et al.NeurIPS 2024 · 16 citations
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 15 citations
- Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement LearningJeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul SungNeurIPS 2024 · 11 citations
Builds on24
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Online Pre-Training for Offline-to-Online Reinforcement LearningYongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong et al.ICML 2025
- From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement LearningLipeng Zu, YU QIAN, Shayok Chakraborty, Xiaonan ZhangICML 2026
- State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement LearningSonglin Li, Wei Xiao, Hao Wu, Xiaodan Zhang et al.AAAI 2026
- SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement LearningJiaheng Feng, Mingxiao Feng, Haolin Song, Wengang Zhou et al.AAAI 2024 · 7 citations
- Adaptive Scaling of Policy Constraints for Offline Reinforcement LearningJing Tan, Xiaorui Li, Chao Yao, Xiaojuan Ban et al.ICLR 2026 · 1 citation
