D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
Haoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan, Jie Fu
摘要
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model's fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang 等NeurIPS 2024 · 被引用 50 次
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang 等NeurIPS 2025 · 被引用 33 次
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of MultilingualityShayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I-Hung Hsu 等ICLR 2026 · 被引用 20 次
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu 等ICML 2026 · 被引用 9 次
- Olmix: A Framework for Data Mixing Throughout LM DevelopmentMayee Chen, Tyler Murray, David Heineman, Matt Jordan 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper21
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng 等AAAI 2020 · 被引用 885 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos 等ICLR 2024 · 被引用 433 次
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch 等ICML 2022 · 被引用 266 次
相关 Paper
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li 等ICML 2025
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan 等ICLR 2025
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-trainingLei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen 等ACL 2026
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language ModelsSamira Abnar, Harshay Shah, Dan Busbridge, Alaaeldin El-Nouby 等ICML 2025
