Transformers without Normalization
Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu
Abstract
Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation DyT(x) = tanh(ωx), as a dropin replacement for normalization layers in Transformers. DyT is inspired by the observation that layer normalization in Transformers often produces tanh-like, S-shaped input-output mappings. By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning. We validate the effectiveness of Transformers with DyT across diverse settings, from recognition to generation, supervised to self-supervised learning, and computer vision to language models. These findings challenge the conventional understanding that normalization layers are indispensable in modern neural networks, and offer new insights into their role in deep networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear MappingHaonan Dong, Wenhao Zhu, Guojie Song, Liang WangNeurIPS 2025 · 31 citations
- Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM KernelsMaximilian Beck, Korbinian Pöppel, Phillip Lippe, Sepp HochreiterNeurIPS 2025 · 17 citations
- Stronger Normalization-Free TransformersMingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun et al.CVPR 2026 · 16 citations
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 13 citations
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 7 citations
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Is normalization indispensable for training deep neural network?Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue et al.NeurIPS 2020 · 70 citations
- Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language ModelsHoyoon Byun, Youngjun Choi, Taero Kim, Sungrae Park et al.ICML 2026 · 2 citations
- SeeDNorm: Self-Rescaled Dynamic NormalizationWenrui Cai, Defa Zhu, Siyuan Qiao, Qingjie Liu et al.ICLR 2026 · 7 citations
- Unified Normalization for Accelerating and Stabilizing TransformersQiming Yang, Kai Zhang, Chaoxiang Lan, Zhi Yang et al.ACM MM 2022 · 1 citation
- Impact of Layer Norm on Memorization and Generalization in TransformersRishi Singhal, Jung-Eun KimNeurIPS 2025 · 4 citations
