Searching for Efficient Transformers for Language Modeling
David R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, Quoc V. Le
摘要
Large Transformer models have been central to recent advances in natural language processing. The training and inference costs of these models, however, have grown rapidly and become prohibitively expensive. Here we aim to reduce the costs of Transformers by searching for a more efficient variant. Compared to previous approaches, our search is performed at a lower level, over the primitives that define a Transformer TensorFlow program. We identify an architecture, named Primer, that has a smaller training cost than the original Transformer and other variants for auto-regressive language modeling. Primer's improvements can be mostly attributed to two simple modifications: squaring ReLU activations and adding a depthwise convolution layer after each Q, K, and V projection in self-attention. Experiments show Primer's gains over Transformer increase as compute scale grows and follow a power law with respect to quality at optimal model sizes. We also verify empirically that Primer can be dropped into different codebases to significantly speed up training without additional tuning. For example, at a 500M parameter size, Primer improves the original T5 architecture on C4 auto-regressive language modeling, reducing the training cost by 4X. Furthermore, the reduced training cost means Primer needs much less compute to reach a target one-shot performance. For instance, in a 1.9B parameter configuration similar to GPT-3 XL, Primer uses 1/3 of the training compute to achieve the same one-shot performance as Transformer. We open source our models and several comparisons in T5 to help with reproducibility. 1 1 https://github.com/google-research/google-research/tree/master/primer 2 We provide details of our primitives search in TensorFlow, but the same approach can also be applied to other deep learning libraries. 35th Conference on Neural Information Processing Systems (NeurIPS 2021), virtual.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun 等NeurIPS 2022 · 被引用 247 次
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik 等NeurIPS 2023 · 被引用 212 次
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- The Efficiency MisnomerMostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer 等ICLR 2022 · 被引用 116 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Evaluating The Search Phase of Neural Architecture SearchKaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat 等ICLR 2020 · 被引用 370 次
相关 Paper
- E.T.: re-thinking self-attention for transformer models on GPUsShiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li 等SC 2021 · 被引用 13 次
- LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language ModelsMojan Javaheripi, Gustavo de Rosa, Subhabrata Mukherjee, Shital Shah 等NeurIPS 2022 · 被引用 27 次
- Primer: Fast Private Transformer Inference on Encrypted DataMengxin Zheng, Qian Lou, Lei JiangDAC 2023 · 被引用 26 次
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King 等NeurIPS 2023 · 被引用 46 次
