Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
Wenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu
Abstract
LLMs are computationally expensive to pre-train due to their large scale. Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones. However, the viability of these model growth methods in efficient LLM pre-training remains underexplored. This work identifies three critical bstacles: (1) lack of comprehensive evaluation, (2) untested viability for scaling, and (3) lack of empirical guidelines. To tackle 1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting. Our findings reveal that a depthwise stacking operator, called , exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines. Motivated by these promising results, we conduct extensive experiments to delve deeper into to address 2 and 3. For 2 (untested scalability), our study shows that is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens. For example, compared to a conventionally trained 7B model using 300B tokens, our model converges to the same loss with 194B tokens, resulting in a 54.6% speedup. We further address 3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for , making it practical in general LLM pre-training. We also provide in-depth discussions and comprehensive ablation studies of . Our code and pre-trained model are available at https://llm-stacking.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fccd65c-694e-4a49-9e8b-933407f00fc9Cited by top-tier papers13
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao et al.ACL 2025 · 6 citations
- When Does Sparsity Mitigate the Curse of Depth in LLMsYao Yao, Xinyuan Song, Sebastian Pokutta, Max Zimmer et al.ICML 2026 · 5 citations
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive LearningQifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang et al.ICML 2026 · 3 citations
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald et al.ICML 2026 · 2 citations
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou et al.ICLR 2026 · 1 citation
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
Related papers
- Masked Structural Growth for 2x Faster Language Model Pre-trainingYiqun Yao, Zheng Zhang, Jing Li, Yequan WangICLR 2024 · 30 citations
- Deep Fusion: Efficient Network Training via Pre-trained InitializationsHanna Mazzawi, Javier Gonzalvo, Michael Wunder, Sammy Jerome et al.ICML 2024 · 4 citations
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi et al.NeurIPS 2024 · 23 citations
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 1 citation
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LNPengxiang Li, Lu Yin, Shiwei LiuICLR 2025
