Midtraining Bridges Pretraining and Posttraining Distributions
Emmy Liu, Graham Neubig, Chenyan Xiong
摘要
Midtraining, the practice of mixing specialized data with more general pretraining data in an intermediate training phase, has become widespread in language model development, yet there is little understanding of what makes it effective. We propose that midtraining functions as distributional bridging by providing better initialization for posttraining. We conduct controlled pretraining experiments, and find that midtraining benefits are largest for domains distant from general pretraining data, such as code and math, and scale with the proximity advantage the midtraining data provides toward the target distribution. In these domains, midtraining consistently outperforms continued pretraining on specialized data alone both in-domain and in terms of mitigating forgetting. We further conduct an investigation on the starting time and mixture weight of midtraining data, using code as a case study, and find that time of introduction and mixture weight interact strongly such that early introduction of specialized data is amenable to high mixture weights, while late introduction requires lower ones. This suggests that late introduction of specialized data outside a plasticity window cannot be compensated for by increasing data mixtures later in training. Beyond midtraining itself, this suggests that distributional transitions between any training phases may benefit from similar bridging strategies. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 被引用 58 次
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim 等ICML 2026 · 被引用 22 次
- Sharpness-Aware Pretraining Mitigates Catastrophic ForgettingIshaan Watts, Catherine Li, Sachin Goyal, Jacob Mitchell Springer 等ICML 2026 · 被引用 6 次
- Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language ModelsShaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen 等ACL 2026 · 被引用 1 次
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMsShane Bergsma, Nolan Simran Dey, Joel HestnessICLR 2026
它引用的顶会 Paper11
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian 等NeurIPS 2022 · 被引用 69 次
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMsFeiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar 等ICLR 2024 · 被引用 39 次
相关 Paper
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-trainingKailai Yang, Xiao Liu, Lei Ji, Hao Li 等ACL 2026 · 被引用 3 次
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni 等ICML 2025
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- Improving Language Plasticity via Pretraining with Active ForgettingYihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani 等NeurIPS 2023 · 被引用 49 次
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan 等ICLR 2025
