Simplifying Transformer Blocks
Bobby He, Thomas Hofmann
摘要
A simple design recipe for deep Transformers is to compose identical building blocks. But standard transformer blocks are far from simple, interweaving attention and MLP sub-blocks with skip connections&normalisation layers in precise arrangements. This complexity leads to brittle architectures, where seemingly minor changes can significantly reduce training speed, or render models untrainable. In this work, we ask to what extent the standard transformer block can be simplified? Combining signal propagation theory and empirical observations, we motivate modifications that allow many block components to be removed with no loss of training speed, including skip connections, projection or value parameters, sequential sub-blocks and normalisation layers. In experiments on both autoregressive decoder-only and BERT encoder-only models, our simplified transformers emulate the per-update training speed and performance of standard transformers, while enjoying 15% faster training throughput, and using 15% fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung 等ICML 2024 · 被引用 16 次
- Stronger Normalization-Free TransformersMingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun 等CVPR 2026 · 被引用 16 次
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper25
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
相关 Paper
- Attention-Only Transformers via Unrolled Subspace DenoisingPeng Wang, Yifu Lu, Yaodong Yu, Druv Pai 等ICML 2025
- An Efficient Transformer Decoder with Compressed Sub-layersYanyang Li, Ye Lin, Tong Xiao, Jingbo ZhuAAAI 2021 · 被引用 32 次
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 被引用 912 次
- Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal PropagationBobby He, James Martens, Guodong Zhang, Aleksandar Botev 等ICLR 2023 · 被引用 5 次
- What Layers When: Learning to Skip Compute in LLMs with Residual GatesFilipe Laitenberger, Dawid Jan Kopiczko, Cees G. M. Snoek, Yuki M. AsanoICLR 2026 · 被引用 6 次
