Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
Howard Chen, Noam Razin, Karthik Narasimhan, Danqi Chen
Abstract
Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities -- a phenomenon classically known as catastrophic forgetting. In this paper, toward identifying guidelines for mitigating this phenomenon, we systematically compare the forgetting patterns of two widely adopted post-training methods: supervised fine-tuning (SFT) and reinforcement learning (RL). Our experiments reveal a consistent trend across LM families (Llama, Qwen) and tasks (instruction following, general knowledge, and arithmetic reasoning): RL leads to less forgetting than SFT while achieving comparable or higher target task performance. To investigate the cause for this difference, we consider a simplified setting in which the LM is modeled as a mixture of two distributions, one corresponding to prior knowledge and the other to the target task. We identify that the mode-seeking nature of RL, which stems from its use of on-policy data, enables keeping prior knowledge intact when learning the target task. We then verify this insight by demonstrating that the use on-policy data underlies the robustness of RL to forgetting in practical settings, as opposed to other algorithmic choices such as the KL regularization or advantage estimation. Lastly, as a practical implication, our results highlight the potential of mitigating forgetting using approximately on-policy data, which can be substantially more efficient to obtain than fully on-policy data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42e1f806-26e9-4b3d-bc2f-824dcb9f12afCited by top-tier papers6
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma et al.ICML 2026 · 46 citations
- Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan et al.ACL 2026 · 5 citations
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang et al.ICML 2026 · 3 citations
- Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors PersistentXingyi Zhao, Tian Xie, Xiaojun Qi, Depeng Xu et al.ICML 2026
- Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document ParsingJun-Peng Jiang, An-Yang Ji, Shiyin Lu, Guodong Zheng et al.ICML 2026
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
Related papers
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 176 citations
- TMS: Trajectory-Mixed Supervision for On-Policy Self DistillationRana Khan, Zijie Liu, Zhen Tan, Charles Fleming et al.ICML 2026
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data PerspectiveZhihao Zhang, Qiaole Dong, Qi Zhang, Enyu Zhou et al.ICLR 2026 · 15 citations
- Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language ModelsChao Xue, Yao Wang, Mengqiao Liu, Di Liang et al.ACL 2026 · 5 citations
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu et al.ICML 2026 · 102 citations
