Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
Syeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, Yejin Choi, Bryan Catanzaro
Abstract
The prevailing paradigm for enhancing the reasoning abilities of Large Language Models (LLMs) revolves around post-training on high-quality, reasoning-intensive data. While emerging literature suggests that reasoning data is increasingly incorporated also during the mid-training stage---a practice that is relatively more proprietary and less openly characterized---the role of such data in pretraining remains unclear. In particular, due to the opaqueness of pretraining corpora in most frontier models, the effect of reasoning data introduced at different phases of pre- and/or post-training is relatively less reported in the scientific literature. This raises several important but unsettled questions: Is adding reasoning data earlier during pre-training any better than introducing it during post-training, when the token counts are controlled? Could earlier inclusion risk overfitting and harm generalization, or instead establish durable foundations that later fine-tuning cannot recover? To address these questions, we conduct the first systematic study of how reasoning data—varying in scale, diversity, and quality—affects LLM performance when introduced at different stages of training. Our findings reveal that front-loading reasoning data into pretraining is critical (19% average gain), establishing foundational capabilities that cannot be fully replicated by later-stage SFT, even with more data. We uncover an asymmetric principle for optimal data allocation: pretraining benefits most from broad diversity in reasoning patterns (11% average gain), while SFT is more sensitive to data quality (15% average gain with high quality data). Furthermore, we show that high-quality pretraining data has latent effects, activated only after SFT, and that naively scaling SFT data can be detrimental, washing away the benefits of early reasoning injection. Collectively, our results challenge the conventional separation of language modeling and reasoning, providing a principled guide for strategically allocating data across the entire training pipeline to build more capable models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 58 citations
- Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning ModelsAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICML 2026 · 4 citations
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
- Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language ModelsShaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen et al.ACL 2026 · 1 citation
- Reasoning Quality Emerges Early: Data Curation for Reasoning ModelsHongyi Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato et al.ICML 2026
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- At Which Training Stage Does Code Data Help LLMs Reasoning?Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang et al.ICLR 2024 · 106 citations
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in ReasoningWang Yang, Zirui Liu, Hongye Jin, Qingyu Yin et al.NeurIPS 2025 · 7 citations
- Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models ReasoningXinlu Zhang, Zhiyu Zoey Chen, Xi Ye, Xianjun Yang et al.AAAI 2025 · 40 citations
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math ProblemsTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 2 citations
- Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training StagesZui Chen, Tianqiao Liu, Tongqing, Mi Tian et al.ICLR 2025
