Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
Wenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu
摘要
LLMs are computationally expensive to pre-train due to their large scale. Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones. However, the viability of these model growth methods in efficient LLM pre-training remains underexplored. This work identifies three critical bstacles: (1) lack of comprehensive evaluation, (2) untested viability for scaling, and (3) lack of empirical guidelines. To tackle 1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting. Our findings reveal that a depthwise stacking operator, called , exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines. Motivated by these promising results, we conduct extensive experiments to delve deeper into to address 2 and 3. For 2 (untested scalability), our study shows that is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens. For example, compared to a conventionally trained 7B model using 300B tokens, our model converges to the same loss with 194B tokens, resulting in a 54.6% speedup. We further address 3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for , making it practical in general LLM pre-training. We also provide in-depth discussions and comprehensive ablation studies of . Our code and pre-trained model are available at https://llm-stacking.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao 等ACL 2025 · 被引用 6 次
- When Does Sparsity Mitigate the Curse of Depth in LLMsYao Yao, Xinyuan Song, Sebastian Pokutta, Max Zimmer 等ICML 2026 · 被引用 5 次
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive LearningQifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang 等ICML 2026 · 被引用 3 次
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald 等ICML 2026 · 被引用 2 次
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- Masked Structural Growth for 2x Faster Language Model Pre-trainingYiqun Yao, Zheng Zhang, Jing Li, Yequan WangICLR 2024 · 被引用 30 次
- Deep Fusion: Efficient Network Training via Pre-trained InitializationsHanna Mazzawi, Javier Gonzalvo, Michael Wunder, Sammy Jerome 等ICML 2024 · 被引用 4 次
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi 等NeurIPS 2024 · 被引用 23 次
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 被引用 1 次
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LNPengxiang Li, Lu Yin, Shiwei LiuICLR 2025
