CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
Jiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao, Fei Tan
摘要
Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT) enhances LLM capabilities by imbuing new domain-specific or proprietary knowledge while replaying general corpus to prevent catastrophic forgetting. The data mixture ratio of general corpus and domain-specific corpus, however, has been chosen heuristically, leading to sub-optimal training efficiency in practice. In this context, we attempt to re-visit the scaling behavior of LLMs under the hood of CPT, and discover a power-law relationship between loss, mixture ratio, and training tokens scale. We formalize the trade-off between general and domain-specific capabilities, leading to a well-defined Critical Mixture Ratio (CMR) of general and domain data. By striking the balance, CMR maintains the model's general ability and achieves the desired domain transfer, ensuring the highest utilization of available resources. Considering the balance between efficiency and effectiveness, CMR can be regarded as the optimal mixture ratio. Through extensive experiments, we ascertain the predictability of CMR, propose CMR scaling law and have substantiated its generalization. These findings offer practical guidelines for optimizing LLM training in specialized domains, ensuring both general and domain-specific performance while efficiently managing training resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Scaling Laws for Optimal Data MixturesMustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier 等NeurIPS 2025 · 被引用 54 次
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced PolicyShaoxiong Zhan, Yanlin Lai, Ziyu Lu, Dahua Lin 等AAAI 2026 · 被引用 17 次
- MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model MergingJiapeng Wang, Changxin Tian, Kunlong Chen, ziqi liu 等ICML 2026 · 被引用 6 次
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and RepetitionWeidong Zhou, Fengze Liu, LIU, Ping Guo 等ICML 2026 · 被引用 1 次
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuningYiming Du, Yifan Xiang, Bin Liang, Dahua Lin 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Understanding Emergent Abilities of Language Models from the Loss PerspectiveZhengxiao Du, Aohan Zeng, Yuxiao Dong, Jie TangNeurIPS 2024 · 被引用 113 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang 等NeurIPS 2024 · 被引用 47 次
相关 Paper
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song 等ICLR 2025
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li 等ICML 2025
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan 等ICLR 2025
- Breaking Language Barriers: Cross-Lingual Continual Pre-Training at ScaleWenzhen Zheng, Wenbo Pan, Xu Xu, Libo Qin 等EMNLP 2024 · 被引用 3 次
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-trainingLei Liu, Hao Zhu, Xiaoyan Yang, Yue Shen 等ACL 2026
