Scaling with Collapse: Efficient and Predictable Training of LLM Families
Shane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad, Gurpreet Gosal, Joel Hestness
Abstract
Effective LLM training depends on predictable scaling of key quantities-such as final loss and optimal hyperparameters-with model and dataset size. Qiu et al. (2025) recently showed that this predictability can extend beyond scalars: whole training loss curves can collapse onto a universal trajectory after a simple normalization. What remains unclear is whether this phenomenon persists for LLM families trained under practical scaling recipes, where width, depth, learning rate, batch size, and weight decay are scaled jointly. We show that it does: loss curves collapse across scales precisely when optimization hyperparameters are set optimally for the given data budget, in accordance with recent empirical scaling laws. Collapse therefore emerges as a signature of compute-efficient training. We demonstrate two applications at scale: (1) deviation-from-collapse provides a sensitive, early diagnostic of training pathologies, and (2) predictability of collapsed curves enables early stopping in large-scale hyperparameter tuning. Finally, we train a competitive LLM family, Celerity, using these insights, establishing collapse as an effective tool for developing efficient LLMs. All experiments were run on Cerebras CS-3 systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 6 citations
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang et al.ICML 2026
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMsShane Bergsma, Nolan Simran Dey, Joel HestnessICLR 2026
Builds on38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
Related papers
- Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural NetworksShikai Qiu, Lechao Xiao, Andrew Gordon Wilson, Jeffrey Pennington et al.ICML 2025
- Scaling Law with Learning Rate AnnealingHowe Tissue, Venus Wang, Lu WangNeurIPS 2025 · 33 citations
- A Hitchhiker's Guide to Scaling Law EstimationLeshem Choshen, Yang Zhang, Jacob AndreasICML 2025
- Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingShane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray et al.NeurIPS 2025 · 44 citations
- A Multi-Power Law for Loss Curve Prediction Across Learning Rate SchedulesKairong Luo, Haodong Wen, Shengding Hu, Zhenbo Sun et al.ICLR 2025
