MA-BERT: Towards Matrix Arithmetic-only BERT Inference by Eliminating Complex Non-Linear Functions
Neo Wei Ming, Zhehui Wang, Cheng Liu, Rick Siow Mong Goh, Tao Luo
Abstract
Due to their superior results, Transformer-based models such as BERT have become de facto standards in many Natural Language Processing (NLP) applications. However, the intensive use of complex non-linear functions within the Transformer architecture impairs its computing efficiency and complicates corresponding accelerator designs, because non-linear functions are generally computation-intensive and require special hardware support. In light of this, we propose MA-BERT, which allows matrix arithmetic-only operations in Transformer-based NLP models and achieves efficient inference with negligible accuracy loss. Specifically, we propose four correlated techniques that include approximating softmax with a two-layer neural network, replacing GELU with ReLU, fusing normalization layers with adjacent linear layers, and leveraging knowledge transfer from baseline models. Through these techniques, we are able to eliminate the major non-linear functions in Transformer-based models and obtain MA-BERT with only matrix arithmetic and trivial ReLU operations without compromising on accuracy. With mainly regular matrix arithmetic operations, MA-BERT enables hardware-friendly processing on various computing engines, including CPUs and GPUs. Our experimental results show that MA-BERT achieves up to 27% and 41% reduction in inference time on CPU and GPU, respectively, with comparable accuracy on many downstream tasks compared to the baseline BERT models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cde5ff79-6376-4c7c-92db-d1d2d4901fb1Cited by top-tier papers1
Ask how each one uses itRelated papers
- NN-LUT: neural approximation of non-linear operations for efficient transformer inferenceJoonsang Yu, Junki Park, Seongmin Park, Minsoo Kim et al.DAC 2022 · 63 citations
- I-BERT: Integer-only BERT QuantizationSehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney et al.ICML 2021 · 439 citations
- Softermax: Hardware/Software Co-Design of an Efficient Softmax for TransformersJacob R. Stevens, Rangharajan Venkatesan, Steve Dai, Brucek Khailany et al.DAC 2021 · 143 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-TuningJanghyeon Kim, Janghwan Lee, Jungwook Choi, JeongHo Han et al.DAC 2023 · 9 citations
