Do Transformer Modifications Transfer Across Implementations and Applications?
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li
Abstract
The research community has proposed copious modifications to the Transformer architecture since it was introduced over three years ago, relatively few of which have seen widespread adoption. In this paper, we comprehensively evaluate many of these modifications in a shared experimental setting that covers most of the common uses of the Transformer in natural language processing. Surprisingly, we find that most modifications do not meaningfully improve performance. Furthermore, most of the Transformer variants we found beneficial were either developed in the same codebase that we used or are relatively minor changes. We conjecture that performance improvements may strongly depend on implementation details and correspondingly make some recommendations for improving the generality of experimental results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b54c801-8a42-4f2e-9b18-71fc1381341dCited by top-tier papers11
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language ModelsIman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C. del Mundo et al.ICLR 2024 · 109 citations
- What Do NLP Researchers Believe? Results of the NLP Community MetasurveyJulian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller et al.ACL 2023 · 16 citations
- Sparse Upcycling: Training Mixture-of-Experts from Dense CheckpointsAran Komatsuzaki, Joan Puigcerver, James Lee-Thorp, Carlos Riquelme Ruiz et al.ICLR 2023 · 12 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 273 citations
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson et al.ICLR 2021 · 11 citations
Related papers
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 28 citations
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 87 citations
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 49 citations
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 53 citations
- IOT: Instance-wise Layer Reordering for Transformer StructuresJinhua Zhu, Lijun Wu, Yingce Xia, Shufang Xie et al.ICLR 2021 · 8 citations
