SimpleGPT: Improving GPT via A Simple Normalization Strategy
Marco Chen, Xianbiao Qi, Yelin He, Jiaquan Ye, Rong Xiao
摘要
In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. We introduce a simple normalization strategy, termed SimpleNorm, which stabilizes intermediate activation scales by construction. Then, by analyzing the Hessian of the loss with respect to network activations, we theoretically show that SimpleNorm significantly reduces the spectral norm of the Hessian, thereby permitting larger stable learning rates. We validate our theoretical findings through extensive experiments on large GPT models at parameter scales 1B, 1.4B, 7B and 8B. Empirically, SimpleGPT, our SimpleNorm-based network, tolerates learning rates 3×-10× larger than standard convention, consistently demonstrates strong optimization stability, and achieves substantially better performance than well-established baselines. Specifically, when training 7B-scale models for 60K steps, SimpleGPT achieves a training loss that is 0.08 lower than that of LLaMA2 with QKNorm, reducing the loss from 2.290 to 2.208. Our source code will be released at https: //github.com/Ocram7/SimpleGPT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- Stronger Normalization-Free TransformersMingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun 等CVPR 2026 · 被引用 16 次
- LipsFormer: Introducing Lipschitz Continuity to Vision TransformersXianbiao Qi, Jianan Wang, Yihao Chen, Yukai Shi 等ICLR 2023 · 被引用 4 次
- DNT: a Deeply Normalized Transformer that can be trained by Momentum SGDXianbiao Qi, Marco Chen, Wenjie Xiao, Jiaquan Ye 等ICLR 2026 · 被引用 1 次
相关 Paper
- Taming Transformer Without Using Learning Rate WarmupXianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li 等ICLR 2025
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 被引用 1 次
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation ScalingTianhao Chen, Xin Xu, Zijing Liu, Pengxiang Li 等NeurIPS 2025 · 被引用 2 次
- What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian AnalysisWeronika Ormaniec, Felix Dangel, Sidak Pal SinghICLR 2025
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang 等ICLR 2024 · 被引用 264 次
