Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
Kosuke Nishida, Kyosuke Nishida, Kuniko Saito
摘要
Loss spikes, a phenomenon in which the loss value diverges suddenly, is a fundamental issue in the pre-training of large language models. This paper supposes that the non-uniformity of the norm of the parameters is one of the causes of loss spikes. Here, in training of neural networks, the scale of the gradients is required to be kept constant throughout the layers to avoid the vanishing and exploding gradients problem. However, to meet these requirements in the Transformer model, the norm of the model parameters must be non-uniform, and thus, parameters whose norm is smaller are more sensitive to the parameter update. To address this issue, we propose a novel technique, weight scaling as reparameterization (WeSaR). WeSaR introduces a gate parameter per parameter matrix and adjusts it to the value satisfying the requirements. Because of the gate parameter, WeSaR sets the norm of the original parameters uniformly, which results in stable training. Experimental results with the Transformer decoders consisting of 130 million, 1.3 billion, and 13 billion parameters showed that WeSaR stabilizes and accelerates training and that it outperformed compared methods including popular initialization methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient ClippingGuoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng 等ICML 2026 · 被引用 3 次
- YuLan-Mini: Pushing the Limits of Open Data-efficient Language ModelYiwen Hu, Huatong Song, Jie Chen, Jia Deng 等ACL 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang 等ICLR 2023 · 被引用 295 次
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 被引用 181 次
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 等NeurIPS 2022 · 被引用 161 次
相关 Paper
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation ScalingTianhao Chen, Xin Xu, Zijing Liu, Pengxiang Li 等NeurIPS 2025 · 被引用 2 次
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung 等ICML 2024 · 被引用 16 次
- SpanNorm: Reconciling Training Stability and Performance in Deep TransformersChao Wang, Bei Li, Jiaqi Zhang, Xinyu Liu 等ICML 2026
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 被引用 214 次
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM TrainingTianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu 等ICLR 2025
