What Makes Value Learning Efficient in Residual Reinforcement Learning?
Guozheng Ma, Lu Li, Haoyu Wang, Zixuan Liu, Pierre-Luc Bacon, Dacheng Tao
Abstract
Residual reinforcement learning (RL) enables stable online refinement of expressive pretrained policies by freezing the base and learning only bounded corrections. However, value learning in residual RL poses unique challenges that remain poorly understood. In this work, we identify two key bottlenecks: cold start pathology, where the critic lacks knowledge of the value landscape around the base policy, and structural scale mismatch, where the residual contribution is dwarfed by the base action. Through systematic investigation, we uncover the mechanisms underlying these bottlenecks, revealing that simple yet principled solutions suffice: base-policy transitions serve as an essential value anchor for implicit warmup, and critic normalization effectively restores representation sensitivity for discerning value differences. Based on these insights, we propose DAWN (Data-Anchored Warmup and Normalization), a minimal approach targeting efficient value learning in residual RL. By addressing these bottlenecks, DAWN demonstrates substantial efficiency gains across diverse benchmarks, policy architectures, and observation modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19c6b6bb-8328-4b19-9dd8-0f8642e27149Builds on14
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 470 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
Related papers
- Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline DataZhiyuan Zhou, Andy Peng, Qiyang Li, Sergey Levine et al.ICLR 2025
- Cerebellar-Inspired Residual Control for Fault Recovery: From Inference-Time Adaptation to Structural ConsolidationNethmi Jayasinghe, Diana Gontero, Spencer Brown, Vinod Sangwan et al.ICML 2026 · 1 citation
- Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline DataJeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong et al.ICML 2025
- FANS: A Flatness-Aware Network Structure for Generalization in Offline Reinforcement LearningDa Wang, Yi Ma, Ting Guo, Hongyao Tang et al.NeurIPS 2025
- SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement LearningJiaheng Feng, Mingxiao Feng, Haolin Song, Wengang Zhou et al.AAAI 2024 · 7 citations
