Masked Structural Growth for 2x Faster Language Model Pre-training
Yiqun Yao, Zheng Zhang, Jing Li, Yequan Wang
Abstract
Accelerating large language model pre-training is a critical issue in present research. In this paper, we focus on speeding up pre-training by progressively growing from a small Transformer structure to a large one. There are two main research problems associated with progressive growth: determining the optimal growth schedule, and designing efficient growth operators. In terms of growth schedule, the impact of each single dimension on a schedule's efficiency is under-explored by existing work. Regarding the growth operators, existing methods rely on the initialization of new weights to inherit knowledge, and achieve only non-strict function preservation, limiting further improvements on training dynamics. To address these issues, we propose Masked Structural Growth (MSG), including (i) growth schedules involving all possible dimensions and (ii) strictly function-preserving growth operators that is independent of the initialization of new weights. Experiments show that MSG is significantly faster than related work: we achieve up to 2.2x speedup in pre-training different types of language models while maintaining comparable or better downstream performances. Code is publicly available at https://github.com/cofe-ai/MSG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e45b35b-eb58-41ef-bf96-b38036d016d3Cited by top-tier papers13
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi et al.NeurIPS 2024 · 23 citations
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao et al.ACL 2025 · 6 citations
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive LearningQifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang et al.ICML 2026 · 3 citations
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald et al.ICML 2026 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 236 citations
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin et al.ICML 2020 · 184 citations
Related papers
- Staged Training for Transformer Language ModelsSheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge et al.ICML 2022 · 52 citations
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard et al.ICLR 2023 · 13 citations
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang et al.ACL 2022
- LOIRE: LifelOng learning on Incremental data via pre-trained language model gRowth EfficientlyXue Han, Yitong Wang, Junlan Feng, Wenchun Gao et al.ICLR 2025
- Building on Efficient Foundations: Effective Training of LLMs with Structured Feedforward LayersXiuying Wei, Skander Moalla, Razvan Pascanu, Caglar GulcehreNeurIPS 2024 · 17 citations
