Context Distillation Retains Post-Training Capabilities in Continually Trained LMs
Shankar Padmanabhan, Mustafa Omer Gul, Tanya Goyal
摘要
Post-training endows pretrained LLMs with a variety of desirable skills, such as instruction-following, reasoning, and others. However, these post-trained LLMs only encode knowledge up to a cut-off date, necessitating continual adaptation. Unfortunately, existing solutions cannot effectively learn new knowledge from adaptation document corpora and simultaneously mitigate the forgetting of earlier learned capabilities. To address this, we introduce Distillation via Split Contexts (DiSC), a simple context-distillation based approach for continual knowledge adaptation. DiSC derives student and teacher distributions by conditioning on distinct segments of the training example and minimizes the KL divergence between them for the common tokens. This insight allows us to efficiently apply context-distillation without requiring explicit generation steps during training. We run experiments on three post-trained models and two adaptation domains. Compared to prior finetuning and distillation methods for continual adaptation, DiSC consistently reports the best trade-off between learning new knowledge and mitigating forgetting of previously learned skills like instruction-following and reasoning, or factual knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 被引用 176 次
- Understanding Catastrophic Forgetting in Language Models via Implicit InferenceSuhas Kotha, Jacob Mitchell Springer, Aditi RaghunathanICLR 2024 · 被引用 131 次
- Mass-Editing Memory in a TransformerKevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov 等ICLR 2023 · 被引用 52 次
相关 Paper
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song 等ICLR 2025
- Adversarial Latent Embedding Repair for LLM Continual LearningXilin Xia, Xialiang Tong, Jie Wang, Chi Ma 等ICML 2026
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen 等AAAI 2023 · 被引用 4 次
- LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration DistillationZican Dong, Junyi Li, Jinhao Jiang, Mingyu Xu 等ACL 2025
- Skill Neologisms: Towards Skill-based Continual LearningAntonin Berthon, Nicolás Astorga, Mihaela van der SchaarICML 2026 · 被引用 1 次
