A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning
Yinmin Zhang, Jie Liu, Chuming Li, Yazhe Niu, Yaodong Yang, Yu Liu, Wanli Ouyang
Abstract
Offline-to-online Reinforcement Learning (O2O RL) aims to improve the performance of offline pretrained policy using only a few online samples. Built on offline RL algorithms, most O2O methods focus on the balance between RL objective and pessimism, or the utilization of offline and online samples. In this paper, from a novel perspective, we systematically study the challenges that remain in O2O RL and identify that the reason behind the slow improvement of the performance and the instability of online finetuning lies in the inaccurate Q-value estimation inherited from offline pretraining. Specifically, we demonstrate that the estimation bias and the inaccurate rank of Q-value cause a misleading signal for the policy update, making the standard offline RL algorithms, such as CQL and TD3-BC, ineffective in the online finetuning. Based on this observation, we address the problem of Q-value estimation by two techniques: (1) perturbed value update and (2) increased frequency of Q-value updates. The first technique smooths out biased Q-value estimation with sharp peaks, preventing early-stage policy exploitation of sub-optimal actions. The second one alleviates the estimation bias inherited from offline pretraining by accelerating learning. Extensive experiments on the MuJoco and Adroit environments demonstrate that the proposed method, named SO2, significantly alleviates Q-value estimation issues, and consistently improves the performance against the state-of-the-art methods by up to 83.1%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd6c3ffe-5c72-4e9c-851c-34c1b9041ab6Cited by top-tier papers10
- CrossLight: Offline-to-Online Reinforcement Learning for Cross-City Traffic Signal ControlQian Sun, Rui Zha, Le Zhang, Jingbo Zhou et al.KDD 2024 · 9 citations
- Flow Matching with Injected Noise for Offline-to-Online Reinforcement LearningYongjae Shin, Jongseong Chae, Jongeui Park, Youngchul SungICLR 2026 · 1 citation
- Online Pre-Training for Offline-to-Online Reinforcement LearningYongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong et al.ICML 2025
- Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline DataZhiyuan Zhou, Andy Peng, Qiyang Li, Sergey Levine et al.ICLR 2025
- SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online TransferNathan S. de Lara, Florian ShkurtiICML 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 15 citations
- Tackling Heavy-Tailed Q-Value Bias in Offline-to-Online Reinforcement Learning with Laplace-Robust ModelingRuibo Guo, Lei Liu, Rui Yang, Junjie Shen et al.ICLR 2026
- In-Context Compositional Q-Learning for Offline Reinforcement LearningQiushui Xu, Yuhao Huang, Yushu Jiang, Wenliang Zheng et al.ICLR 2026
- Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion GenerationXiao Huang, Xu Liu, Enze Zhang, Tong Yu et al.ICML 2025
- State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement LearningSonglin Li, Wei Xiao, Hao Wu, Xiaodan Zhang et al.AAAI 2026
