Learning in Compact Spaces with Approximately Normalized Transformer
Jörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev, Frank Hutter, Michael Hefenbrock
摘要
The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the residual stream. A common solution is to apply regularization and normalization techniques that usually require tuning additional hyperparameters. An alternative is to force all parameters and representations to lie on a hypersphere. This removes the need for regularization and increases convergence speed, but comes with additional costs. In this work, we propose a more holistic, approximate normalization via simple scalar multiplications motivated by the tight concentration of the norms of high-dimensional random vectors. Additionally, instead of applying strict normalization for the parameters, we constrain their norms. These modifications remove the need for weight decay and learning rate warm-up as well, but do not increase the total number of normalization layers. Our experiments with transformer architectures show up to 40% faster convergence compared to GPT models with QK normalization, with only 3% additional runtime cost. When deriving scaling laws, we found that our method enables training with larger batch sizes while preserving the favorable scaling characteristics of classic GPT architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
相关 Paper
- nGPT: Normalized Transformer with Representation Learning on the HypersphereIlya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris GinsburgICLR 2025
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 被引用 1 次
- Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three RegimesMaxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. VetrovNeurIPS 2022 · 被引用 25 次
- A Stable Whitening Optimizer for Efficient Neural Network TrainingKevin Frans, Sergey Levine, Pieter AbbeelNeurIPS 2025 · 被引用 18 次
- Initialization of Large Language Models via Reparameterization to Mitigate Loss SpikesKosuke Nishida, Kyosuke Nishida, Kuniko SaitoEMNLP 2024 · 被引用 2 次
