Searching for Efficient Transformers for Language Modeling
David R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, Quoc V. Le
Abstract
Large Transformer models have been central to recent advances in natural language processing. The training and inference costs of these models, however, have grown rapidly and become prohibitively expensive. Here we aim to reduce the costs of Transformers by searching for a more efficient variant. Compared to previous approaches, our search is performed at a lower level, over the primitives that define a Transformer TensorFlow program. We identify an architecture, named Primer, that has a smaller training cost than the original Transformer and other variants for auto-regressive language modeling. Primer's improvements can be mostly attributed to two simple modifications: squaring ReLU activations and adding a depthwise convolution layer after each Q, K, and V projection in self-attention. Experiments show Primer's gains over Transformer increase as compute scale grows and follow a power law with respect to quality at optimal model sizes. We also verify empirically that Primer can be dropped into different codebases to significantly speed up training without additional tuning. For example, at a 500M parameter size, Primer improves the original T5 architecture on C4 auto-regressive language modeling, reducing the training cost by 4X. Furthermore, the reduced training cost means Primer needs much less compute to reach a target one-shot performance. For instance, in a 1.9B parameter configuration similar to GPT-3 XL, Primer uses 1/3 of the training compute to achieve the same one-shot performance as Transformer. We open source our models and several comparisons in T5 to help with reproducibility. 1 1 https://github.com/google-research/google-research/tree/master/primer 2 We provide details of our primitives search in TensorFlow, but the same approach can also be applied to other deep learning libraries. 35th Conference on Neural Information Processing Systems (NeurIPS 2021), virtual.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3556b4ab-c4e3-441c-97ae-43bfe1cf62b4Cited by top-tier papers43
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun et al.NeurIPS 2022 · 247 citations
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik et al.NeurIPS 2023 · 212 citations
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi et al.CVPR 2024 · 137 citations
- The Efficiency MisnomerMostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer et al.ICLR 2022 · 116 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Evaluating The Search Phase of Neural Architecture SearchKaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat et al.ICLR 2020 · 370 citations
Related papers
- E.T.: re-thinking self-attention for transformer models on GPUsShiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li et al.SC 2021 · 13 citations
- LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language ModelsMojan Javaheripi, Gustavo de Rosa, Subhabrata Mukherjee, Shital Shah et al.NeurIPS 2022 · 27 citations
- Primer: Fast Private Transformer Inference on Encrypted DataMengxin Zheng, Qian Lou, Lei JiangDAC 2023 · 26 citations
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King et al.NeurIPS 2023 · 46 citations
