Scaling depth capacity via zero/one-layer model expansion
Zhiqi Bu
Abstract
Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost. To efficiently train models at scale, progressive training (also known as model expansion) scales up model capacity during training and significantly reduces computation with little performance degradation. In this work, we study the depth expansion of large-scale models through the lens of optimization theory and feature learning, offering insights on the initialization of new layers, hyperparameter transfer, learning rate schedule, and timing of model expansion. Specifically, we propose zero/one-layer progressive training to achieve an optimal tradeoff between computation and loss, with a comprehensive ablations on our expansion strategy. For example, zero/one-layer progressive training on GPT2 can save compute, or equivalently achieve an acceleration, while attaining a loss comparable to that of a fully trained 60-layer model with 7B parameters, thus demonstrating a mixing behavior in terms of loss. Furthermore, scaling laws on LLAMA3 and DeepSeekV3 models show a improvement in compute efficiency, with an increasing advantage at larger scales.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe9db789-b504-4b8d-9cf6-c4f775ec7742Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal et al.NeurIPS 2024 · 168 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- Staged Training for Transformer Language ModelsSheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge et al.ICML 2022 · 52 citations
Related papers
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao et al.ACL 2025 · 6 citations
- Curriculum-Guided Layer Scaling for Language Model PretrainingKaranpartap Singh, Neil Band, Ehsan AdeliICML 2026
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li et al.NeurIPS 2025 · 77 citations
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-ExpertsRuizhe Wang, Yucheng Ding, Xiao Liu, Yaoxiang Wang et al.ICML 2026
- The Curse of Depth in Large Language ModelsWenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin et al.NeurIPS 2025 · 62 citations
