Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models
Akhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung, Harshith Goka, Haejun Lee
Abstract
In spite of their huge success, transformer models remain difficult to scale in depth. In this work, we develop a unified signal propagation theory and provide formulae that govern the moments of the forward and backward signal through the transformer model. Our framework can be used to understand and mitigate vanishing/exploding gradients, rank collapse, and instability associated with high attention scores. We also propose DeepScaleLM, an initialization and scaling scheme that conserves unit output/gradient moments throughout the model, enabling the training of very deep models with 1000 layers. We find that transformer models could be much deeper - our deep models with fewer parameters outperform shallow models in Language Modeling, Speech Translation, and Image Classification, across encoder-only, decoder-only and encoder-decoder variants, for both Pre-LN and Post-LN transformers, for multiple datasets and model sizes. These improvements also translate into improved performance on downstream Question Answering tasks and improved robustness for Image Classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b369359-902a-4789-bbb2-abca8a8c7186Cited by top-tier papers9
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li et al.NeurIPS 2025 · 77 citations
- Tropical Attention: Neural Algorithmic Reasoning for Combinatorial AlgorithmsBaran Hashemi, Kurt Pasque, Christopher Teska, Ruriko YoshidaNeurIPS 2025 · 14 citations
- Normalization in Attention DynamicsNikita Karagodin, Shu Ge, Yury Polyanskiy, Philippe RigolletNeurIPS 2025 · 10 citations
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 1 citation
- Peri-LN: Revisiting Normalization Layer in the Transformer ArchitectureJeonghoon Kim, Byeongchan Lee, Cheonbok Park, Yeontaek Oh et al.ICML 2025
Builds on48
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
Related papers
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 15 citations
- SpanNorm: Reconciling Training Stability and Performance in Deep TransformersChao Wang, Bei Li, Jiaqi Zhang, Xinyu Liu et al.ICML 2026
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation ScalingTianhao Chen, Xin Xu, Zijing Liu, Pengxiang Li et al.NeurIPS 2025 · 2 citations
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- Deep Transformers with Latent DepthXian Li, Asa Cooper Stickland, Yuqing Tang, Xiang KongNeurIPS 2020 · 32 citations
