Staged Training for Transformer Language Models
Sheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge, Matthew E. Peters, Iz Beltagy
Abstract
The current standard approach to scaling transformer language models trains each model size from a different random initialization. As an alternative, we consider a staged training setup that begins with a small model and incrementally increases the amount of compute used for training by applying a "growth operator" to increase the model depth and width. By initializing each stage with the output of the previous one, the training process effectively re-uses the compute from prior stages and becomes more efficient. Our growth operators each take as input the entire training state (including model parameters, optimizer state, learning rate schedule, etc.) and output a new training state from which training continues. We identify two important properties of these growth operators, namely that they preserve both the loss and the "training dynamics" after applying the operator. While the losspreserving property has been discussed previously, to the best of our knowledge this work is the first to identify the importance of preserving the training dynamics (the rate of decrease of the loss during training). To find the optimal schedule for stages, we use the scaling laws from (Kaplan et al., 2020) to find a precise schedule that gives the most compute saving by starting a new stage when training efficiency starts decreasing. We empirically validate our growth operators and staged training for autoregressive language models, showing up to 22% compute savings compared to a strong baseline trained from scratch. Our code is available at https://github. com/allenai/staged-training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0a591ea-584e-4933-9edd-e5c511129252Cited by top-tier papers26
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
- No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsJean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini et al.NeurIPS 2023 · 63 citations
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- Masked Structural Growth for 2x Faster Language Model Pre-trainingYiqun Yao, Zheng Zhang, Jing Li, Yequan WangICLR 2024 · 30 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin et al.ICML 2020 · 184 citations
- GradMax: Growing Neural Networks using Gradient InformationUtku Evci, Bart van Merrienboer, Thomas Unterthiner, Fabian Pedregosa et al.ICLR 2022 · 72 citations
- Shallow-to-Deep Training for Neural Machine TranslationBei Li, Ziyang Wang, Hui Liu, Yufan Jiang et al.EMNLP 2020 · 41 citations
- Shortformer: Better Language Modeling using Shorter InputsOfir Press, Noah A. Smith, Mike LewisACL 2021
Related papers
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard et al.ICLR 2023 · 13 citations
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 1 citation
- TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and SearchingCheng Fu, Hanxian Huang, Zixuan Jiang, Yun Ni et al.ICCV 2023 · 5 citations
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang et al.ACL 2022
- Scaling Law with Learning Rate AnnealingHowe Tissue, Venus Wang, Lu WangNeurIPS 2025 · 33 citations
