Simplifying Transformer Blocks
Bobby He, Thomas Hofmann
Abstract
A simple design recipe for deep Transformers is to compose identical building blocks. But standard transformer blocks are far from simple, interweaving attention and MLP sub-blocks with skip connections&normalisation layers in precise arrangements. This complexity leads to brittle architectures, where seemingly minor changes can significantly reduce training speed, or render models untrainable. In this work, we ask to what extent the standard transformer block can be simplified? Combining signal propagation theory and empirical observations, we motivate modifications that allow many block components to be removed with no loss of training speed, including skip connections, projection or value parameters, sequential sub-blocks and normalisation layers. In experiments on both autoregressive decoder-only and BERT encoder-only models, our simplified transformers emulate the per-update training speed and performance of standard transformers, while enjoying 15% faster training throughput, and using 15% fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcf57fba-ffd7-4cf4-9f47-8ac462c6f2fdCited by top-tier papers13
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag et al.NeurIPS 2024 · 27 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
- Stronger Normalization-Free TransformersMingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun et al.CVPR 2026 · 16 citations
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion et al.ICML 2026 · 7 citations
Builds on25
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
Related papers
- Attention-Only Transformers via Unrolled Subspace DenoisingPeng Wang, Yifu Lu, Yaodong Yu, Druv Pai et al.ICML 2025
- An Efficient Transformer Decoder with Compressed Sub-layersYanyang Li, Ye Lin, Tong Xiao, Jingbo ZhuAAAI 2021 · 32 citations
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
- Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal PropagationBobby He, James Martens, Guodong Zhang, Aleksandar Botev et al.ICLR 2023 · 5 citations
- What Layers When: Learning to Skip Compute in LLMs with Residual GatesFilipe Laitenberger, Dawid Jan Kopiczko, Cees G. M. Snoek, Yuki M. AsanoICLR 2026 · 6 citations
