Understanding Catastrophic Forgetting in Language Models via Implicit Inference
Suhas Kotha, Jacob Mitchell Springer, Aditi Raghunathan
Abstract
We lack a systematic understanding of the effects of fine-tuning (via methods such as instruction-tuning or reinforcement learning from human feedback), particularly on tasks outside the narrow fine-tuning distribution. In a simplified scenario, we demonstrate that improving performance on tasks within the fine-tuning data distribution comes at the expense of capabilities on other tasks. We hypothesize that language models implicitly infer the task of the prompt and that fine-tuning skews this inference towards tasks in the fine-tuning distribution. To test this, we propose Conjugate Prompting, which artificially makes the task look farther from the fine-tuning distribution while requiring the same capability, and we find that this recovers some of the pretraining capabilities in our synthetic setup. Since real-world fine-tuning distributions are predominantly English, we apply conjugate prompting to recover pretrained capabilities in LLMs by simply translating the prompts to different languages. This allows us to recover in-context learning abilities lost via instruction tuning, natural reasoning capability lost during code fine-tuning, and, more concerningly, harmful content generation suppressed by safety fine-tuning in chatbots like ChatGPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18cacbbc-4331-49fd-9c53-1a27ba10349bCited by top-tier papers51
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksSamyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick et al.ICLR 2024 · 108 citations
- What Makes and Breaks Safety Fine-tuning? A Mechanistic StudySamyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy et al.NeurIPS 2024 · 62 citations
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsJiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang et al.WWW 2025 · 42 citations
- Multi-Attribute Steering of Language Models via Targeted InterventionDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2025 · 30 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li et al.ACL 2024 · 39 citations
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis et al.ICML 2026
- On the Loss of Context Awareness in General Instruction Fine-tuningYihan Wang, Andrew Bai, Nanyun Peng, Cho-Jui HsiehNeurIPS 2025 · 11 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik et al.ICLR 2024 · 45 citations
