Learning Dynamics in Continual Pre-Training for Large Language Models
Xingjin Wang, Howe Tissue, Lu Wang, Linjing Li, Daniel Dajun Zeng
摘要
Continual Pre-Training (CPT) is a popular and effective method for applying strong foundation models to specific downstream tasks. In this work, we explore the learning dynamics throughout the CPT process for large language models. We specifically focus on how general and downstream domain performance evolves at each training step, with performance measured by validation losses. We observe that the CPT loss curve fundamentally characterizes a transition from an initial pre-training trajectory to a new, domainspecific one, conceptualized as a shift between two hidden loss curves. This transition can be described by decoupling the effects of distribution shift and learning rate annealing. We derive a CPT scaling law that combines these two factors, enabling the prediction of loss at any (continual) training step and across various learning rate schedules. Our formulation presents a comprehensive understanding of several critical factors in CPT, including loss potential, peak learning rate, training steps, and replay ratio. Moreover, our approach can be adapted to optimize training hyper-parameters for different CPT goals, such as balancing general and domain-specific performance. Extensive experiments demonstrate that our scaling law holds across various CPT datasets and hyper-parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma 等ICML 2026 · 被引用 46 次
- TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsRicardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves 等ACL 2026 · 被引用 34 次
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase 等ICLR 2026 · 被引用 13 次
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-trainingLei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen 等ACL 2026
- Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to AlgorithmHaoyu Wang, yifan shang, Zhongxiang Sun, Weijie Yu 等ICML 2026
它引用的顶会 Paper12
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web TextKeiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy BaICLR 2024 · 被引用 140 次
- Lifelong Language Pretraining with Distribution-Specialized ExpertsWuyang Chen, Yanqi Zhou, Nan Du, Yanping Huang 等ICML 2023 · 被引用 85 次
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang 等NeurIPS 2024 · 被引用 47 次
相关 Paper
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- Scaling Law with Learning Rate AnnealingHowe Tissue, Venus Wang, Lu WangNeurIPS 2025 · 被引用 33 次
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
- On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of CodeMartin Weyssow, Xin Zhou, Kisub Kim, David Lo 等FSE 2023 · 被引用 9 次
- Breaking Language Barriers: Cross-Lingual Continual Pre-Training at ScaleWenzhen Zheng, Wenbo Pan, Xu Xu, Libo Qin 等EMNLP 2024 · 被引用 3 次
