Efficient Domain Continual pretraining by Mitigating the Stability Gap
Yiduo Guo, Jie Fu, Huishuai Zhang, Dongyan Zhao
Abstract
Continual pre-training has increasingly become the predominant approach for adapting Large Language Models (LLMs) to new domains. This process involves updating the pre-trained LLM with a corpus from a new domain, resulting in a shift in the training distribution. To study the behavior of LLMs during this shift, we measured the model's performance throughout the continual pre-training process. we observed a temporary performance drop at the beginning, followed by a recovery phase, a phenomenon known as the "stability gap," previously noted in vision models classifying new classes. The substantial performance drop and slow recovery associated with this gap lead to inefficient pre-training for domain performance improvement and the forgetting of general task knowledge. To address this issue and enhance LLM performance within a fixed compute budget, we propose three effective strategies: (1) Continually pre-training the LLM on a subset with a proper size for multiple epochs, resulting in faster performance recovery than pre-training the LLM on a large corpus in a single epoch; (2) Pretraining the LLM only on high-quality sub-corpus, which rapidly boosts domain performance; and (3) Using a data mixture similar to the pre-training data to reduce distribution gap. We conduct various experiments on Llama-family models to validate the effectiveness of our strategies in both medical continual pre-training and instruction tuning. For example, our strategies improve the average medical task performance of the OpenLlama-3B model from 36.2% to 40.7% with only 40% of the original training budget and enhance the average general task performance without causing forgetting. Furthermore, we apply our strategies to continually pre-train and instruction-tune the Llama-3-8B model. The resulting model, Llama-3-Physician, achieves the best medical performance among current open-source models, and performs comparably to or even better than GPT-4 on several medical benchmarks. We release our models at https://huggingface.co/YiDuo1999/ Llama-3-Physician-8B-Instruct . Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b0d39d0-bf45-43c3-ac15-e2aebd953a85Cited by top-tier papers2
- RedSage: A Cybersecurity Generalist LLMNaufal Suryanto, Muzammal Naseer, Pengfei Li, Syed Talal Wasim et al.ICLR 2026 · 2 citations
- Modular Pretraining Enables Access ControlEthan Roland, Murat Cubuktepe, Erick Martinez, Stijn Servaes et al.ICML 2026
Builds on18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Towards Effective and Efficient Continual Pre-training of Large Language ModelsJie Chen, Zhipeng Chen, Jiapeng Wang, Kun Zhou et al.ACL 2025
- Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text SummarizationGunjan Balde, Soumyadeep Roy, Mainack Mondal, Niloy GangulyACL 2026
- LLaMA Pro: Progressive LLaMA with Block ExpansionChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu et al.ACL 2024
- Can Continual Pretraining Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?Niclas Doll, Jasper Schulze Buschhoff, Shalaka Satheesh, Hammam Abdelwahab et al.ACL 2026 · 1 citation
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song et al.ICLR 2025
