RL's Razor: Why Online Reinforcement Learning Forgets Less
Idan Shenfeld, Jyothish Pari, Pulkit Agrawal
摘要
Comparison of fine-tuning models with reinforcement learning (RL) and supervised fine-tuning (SFT) reveals that, despite similar performance at a new task, RL preserves prior knowledge and capabilities significantly better. We find that the degree of forgetting is determined by the distributional shift, measured as the KL-divergence between the fine-tuned and base policy evaluated on the new task. Our analysis reveals that on-policy RL is implicitly biased towards KL-minimal solutions among the many that solve the new task, whereas SFT can converge to distributions arbitrarily far from the base model. We validate these findings through experiments with large language models and robotic foundation models and further provide theoretical justification for why on-policy RL updates lead to a smaller KL change. We term this principle RL's Razor: among all ways to solve a new task, RL prefers those closest in KL to the original model. Our website is available at http://jyopari.github.io/posts/rl_razor .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 被引用 159 次
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld 等ICLR 2026 · 被引用 116 次
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RLWenli Xiao, Haotian Lin, Andy Peng, Haoru Xue 等ICLR 2026 · 被引用 84 次
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim 等NeurIPS 2025 · 被引用 78 次
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma 等ICML 2026 · 被引用 46 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Retaining by Doing: The Role of On-Policy Data in Mitigating ForgettingHoward Chen, Noam Razin, Karthik Narasimhan, Danqi ChenICML 2026
- TMS: Trajectory-Mixed Supervision for On-Policy Self DistillationRana Khan, Zijie Liu, Zhen Tan, Charles Fleming 等ICML 2026
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data PerspectiveZhihao Zhang, Qiaole Dong, Qi Zhang, Enyu Zhou 等ICLR 2026 · 被引用 15 次
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable RewardLong Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu 等ICLR 2026 · 被引用 46 次
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin 等ACL 2026 · 被引用 1 次
