Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva
Abstract
This work aims to understand how scaling improves language models, specifically in terms of training dynamics. We find that language models undergo loss deceleration early in training; an abrupt slowdown in the rate of loss improvement, resulting in piecewise linear behaviour of the loss curve in log-log space. Scaling up the model mitigates this transition by (1) decreasing the loss at which deceleration occurs, and (2) improving the log-log rate of loss improvement after deceleration. We attribute loss deceleration to a type of degenerate training dynamics we term zero-sum learning (ZSL). In ZSL, per-example gradients become systematically opposed, leading to destructive interference in per-example changes in loss. As a result, improving loss on one subset of examples degrades it on another, bottlenecking overall progress. Loss deceleration and ZSL provide new insights into the training dynamics underlying language model scaling laws, and could potentially be targeted directly to improve language models independent of scale. We make our code and artefacts available at: https://github.com/mirandrom/zsl
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79866f21-3de9-4c02-9b35-d470403d6cebCited by top-tier papers3
- SFTMix: Elevating Language Model Instruction Tuning with Mixup RecipeYuxin Xiao, Shujian Zhang, Marzyeh Ghassemi, Wenxuan ZhouACL 2026 · 3 citations
- Convergent World Representations and Divergent TasksCore Francisco ParkICML 2026
- Understanding the Emergence of Seemingly Useless Features in Next-Token PredictorsMark Rofin, Jalal Naghiyev, Michael HahnICLR 2026
Builds on13
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Learning explanations that are hard to varyGiambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele et al.ICLR 2021 · 221 citations
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
Related papers
- What Scales in Cross-Entropy Scaling Law?Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu et al.ICLR 2026 · 1 citation
- Scaling Law with Learning Rate AnnealingHowe Tissue, Venus Wang, Lu WangNeurIPS 2025 · 33 citations
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 6 citations
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling LawsXu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu et al.ICML 2026
- Inverse Depth Scaling From Most Layers Being SimilarYizhou Liu, Sara Kangaslahti, Ziming Liu, Jeff GoreICML 2026 · 4 citations
