Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
Louis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni, Marco Cuturi, Pierre Ablin
摘要
A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from that target domain. Finetuning presents two challenges: (i) if the amount of target data is limited, as in most practical applications, the model will quickly overfit, and (ii) the model will drift away from the original model, forgetting the pretraining data and the generic knowledge that comes with it. Our goal is to derive scaling laws that quantify these two phenomena for various target domains, amounts of available target data, and model scales. We measure the efficiency of injecting pretraining data into the finetuning data mixture to avoid forgetting and mitigate overfitting. A key practical takeaway from our study is that injecting as little as 1% of pretraining data in the finetuning data mixture prevents the model from forgetting the pretraining set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Scaling Laws for Optimal Data MixturesMustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier 等NeurIPS 2025 · 被引用 54 次
- Midtraining Bridges Pretraining and Posttraining DistributionsEmmy Liu, Graham Neubig, Chenyan XiongICML 2026 · 被引用 8 次
- EvoLM: In Search of Lost Training Dynamics for Language Model ReasoningZhenting Qi, Fan Nie, Alexandre Alahi, James Y. Zou 等NeurIPS 2025 · 被引用 3 次
- Mining Useful General Data for Low-Resource Domain AdaptationPingjie Wang, Hongcheng Liu, Yusheng Liao, Ziqing Fan 等ICML 2026 · 被引用 3 次
- Optimal Splitting of Language Models from Mixtures to Specialized DomainsSkyler Seto, Pierre Ablin, Anastasiia Filippova, Jiayuan Ye 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
相关 Paper
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language ModelsKangtao Lv, Haibin Chen, Yujin Yuan, Langming Liu 等EMNLP 2025
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language ModelsSamira Abnar, Harshay Shah, Dan Busbridge, Alaaeldin El-Nouby 等ICML 2025
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language ModelsYuan Li, Zhengzhong Liu, Eric P. XingICML 2025
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan 等ICLR 2025
