Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva
摘要
This work aims to understand how scaling improves language models, specifically in terms of training dynamics. We find that language models undergo loss deceleration early in training; an abrupt slowdown in the rate of loss improvement, resulting in piecewise linear behaviour of the loss curve in log-log space. Scaling up the model mitigates this transition by (1) decreasing the loss at which deceleration occurs, and (2) improving the log-log rate of loss improvement after deceleration. We attribute loss deceleration to a type of degenerate training dynamics we term zero-sum learning (ZSL). In ZSL, per-example gradients become systematically opposed, leading to destructive interference in per-example changes in loss. As a result, improving loss on one subset of examples degrades it on another, bottlenecking overall progress. Loss deceleration and ZSL provide new insights into the training dynamics underlying language model scaling laws, and could potentially be targeted directly to improve language models independent of scale. We make our code and artefacts available at: https://github.com/mirandrom/zsl
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SFTMix: Elevating Language Model Instruction Tuning with Mixup RecipeYuxin Xiao, Shujian Zhang, Marzyeh Ghassemi, Wenxuan ZhouACL 2026 · 被引用 3 次
- Convergent World Representations and Divergent TasksCore Francisco ParkICML 2026
- Understanding the Emergence of Seemingly Useless Features in Next-Token PredictorsMark Rofin, Jalal Naghiyev, Michael HahnICLR 2026
它引用的顶会 Paper13
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Learning explanations that are hard to varyGiambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele 等ICLR 2021 · 被引用 221 次
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
相关 Paper
- What Scales in Cross-Entropy Scaling Law?Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu 等ICLR 2026 · 被引用 1 次
- Scaling Law with Learning Rate AnnealingHowe Tissue, Venus Wang, Lu WangNeurIPS 2025 · 被引用 33 次
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 被引用 6 次
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling LawsXu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu 等ICML 2026
- Inverse Depth Scaling From Most Layers Being SimilarYizhou Liu, Sara Kangaslahti, Ziming Liu, Jeff GoreICML 2026 · 被引用 4 次
