RL's Razor: Why Online Reinforcement Learning Forgets Less
Idan Shenfeld, Jyothish Pari, Pulkit Agrawal
Abstract
Comparison of fine-tuning models with reinforcement learning (RL) and supervised fine-tuning (SFT) reveals that, despite similar performance at a new task, RL preserves prior knowledge and capabilities significantly better. We find that the degree of forgetting is determined by the distributional shift, measured as the KL-divergence between the fine-tuned and base policy evaluated on the new task. Our analysis reveals that on-policy RL is implicitly biased towards KL-minimal solutions among the many that solve the new task, whereas SFT can converge to distributions arbitrarily far from the base model. We validate these findings through experiments with large language models and robotic foundation models and further provide theoretical justification for why on-policy RL updates lead to a smaller KL change. We term this principle RL's Razor: among all ways to solve a new task, RL prefers those closest in KL to the original model. Our website is available at http://jyopari.github.io/posts/rl_razor .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c04bce6-81dc-43af-a9f8-4ee85d234b39Cited by top-tier papers52
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 159 citations
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld et al.ICLR 2026 · 116 citations
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RLWenli Xiao, Haotian Lin, Andy Peng, Haoru Xue et al.ICLR 2026 · 84 citations
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma et al.ICML 2026 · 46 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Retaining by Doing: The Role of On-Policy Data in Mitigating ForgettingHoward Chen, Noam Razin, Karthik Narasimhan, Danqi ChenICML 2026
- TMS: Trajectory-Mixed Supervision for On-Policy Self DistillationRana Khan, Zijie Liu, Zhen Tan, Charles Fleming et al.ICML 2026
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data PerspectiveZhihao Zhang, Qiaole Dong, Qi Zhang, Enyu Zhou et al.ICLR 2026 · 15 citations
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable RewardLong Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu et al.ICLR 2026 · 46 citations
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin et al.ACL 2026 · 1 citation
