What Makes Value Learning Efficient in Residual Reinforcement Learning?
Guozheng Ma, Lu Li, Haoyu Wang, Zixuan Liu, Pierre-Luc Bacon, Dacheng Tao
摘要
Residual reinforcement learning (RL) enables stable online refinement of expressive pretrained policies by freezing the base and learning only bounded corrections. However, value learning in residual RL poses unique challenges that remain poorly understood. In this work, we identify two key bottlenecks: cold start pathology, where the critic lacks knowledge of the value landscape around the base policy, and structural scale mismatch, where the residual contribution is dwarfed by the base action. Through systematic investigation, we uncover the mechanisms underlying these bottlenecks, revealing that simple yet principled solutions suffice: base-policy transitions serve as an essential value anchor for implicit warmup, and critic normalization effectively restores representation sensitivity for discerning value differences. Based on these insights, we propose DAWN (Data-Anchored Warmup and Normalization), a minimal approach targeting efficient value learning in residual RL. By addressing these bottlenecks, DAWN demonstrates substantial efficiency gains across diverse benchmarks, policy architectures, and observation modalities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu 等ICML 2023 · 被引用 158 次
相关 Paper
- Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline DataZhiyuan Zhou, Andy Peng, Qiyang Li, Sergey Levine 等ICLR 2025
- Cerebellar-Inspired Residual Control for Fault Recovery: From Inference-Time Adaptation to Structural ConsolidationNethmi Jayasinghe, Diana Gontero, Spencer Brown, Vinod Sangwan 等ICML 2026 · 被引用 1 次
- Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline DataJeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong 等ICML 2025
- FANS: A Flatness-Aware Network Structure for Generalization in Offline Reinforcement LearningDa Wang, Yi Ma, Ting Guo, Hongyao Tang 等NeurIPS 2025
- SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement LearningJiaheng Feng, Mingxiao Feng, Haolin Song, Wengang Zhou 等AAAI 2024 · 被引用 7 次
