Instruction Pre-Training: Language Models are Supervised Multitask Learners
Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, Furu Wei
Abstract
Unsupervised multitask pre-training has been the critical method behind the recent success of language models (LMs). However, supervised multitask learning still holds significant promise, as scaling it in the post-training stage trends towards better generalization. In this paper, we explore supervised multitask pretraining by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train LMs. The instruction-response pairs are generated by an efficient instruction synthesizer built on open-source models. In our experiments, we synthesize 200M instruction-response pairs covering 40+ task categories to verify the effectiveness of Instruction Pre-Training. In pre-training from scratch, Instruction Pre-Training not only consistently enhances pre-trained base models but also benefits more from further instruction tuning. In continual pre-training, Instruction Pre-Training enables Llama3-8B to be comparable to or even outperform Llama3-70B. Our model, code, and data are available at https://github.com/microsoft/LMOps .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e8816c6-b3ff-43dc-a563-0fcd35352714Cited by top-tier papers24
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training DataSyeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary et al.ICLR 2026 · 27 citations
- Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM CollaborationSukwon Yun, Jie Peng, Pingzhi Li, Wendong Fan et al.ICLR 2026 · 21 citations
- Precise Information Control in Long-Form Text GenerationJacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li et al.NeurIPS 2025 · 8 citations
- Midtraining Bridges Pretraining and Posttraining DistributionsEmmy Liu, Graham Neubig, Chenyan XiongICML 2026 · 8 citations
- DCR: Quantifying Data Contamination in LLMs EvaluationCheng Xu, Nan Yan, Shuhao Guan, Changhong Jin et al.EMNLP 2025 · 7 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu et al.EMNLP 2025
- Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationLinda He, Jue Wang, Maurice Weber, Shang Zhu et al.ICLR 2025
- Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task GeneralizationYuxian Gu, Pei Ke, Xiaoyan Zhu, Minlie HuangEMNLP 2022 · 3 citations
- Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingZhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng et al.ICLR 2025
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 176 citations
