HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, Song Han
Abstract
Transformers are ubiquitous in Natural Language Processing (NLP) tasks, but they are difficult to be deployed on hardware due to the intensive computation. To enable low-latency inference on resource-constrained hardware platforms, we propose to design Hardware-Aware Transformers (HAT) with neural architecture search. We first construct a large design space with arbitrary encoder-decoder attention and heterogeneous layers. Then we train a Super-Transformer that covers all candidates in the design space, and efficiently produces many SubTransformers with weight sharing. Finally, we perform an evolutionary search with a hardware latency constraint to find a specialized SubTransformer dedicated to run fast on the target hardware. Extensive experiments on four machine translation tasks demonstrate that HAT can discover efficient models for different hardware (CPU, GPU, IoT device). When running WMT'14 translation task on Raspberry Pi-4, HAT can achieve 3× speedup, 3.7× smaller size over baseline Transformer; 2.7× speedup, 3.6× smaller size over Evolved Transformer with 12,041× less search cost and no performance loss. HAT is open-sourced.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9f57220-b94b-44dc-910d-224f2c594a33Cited by top-tier papers60
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
- Representing Long-Range Context for Graph Neural Networks with Global AttentionZhanghao Wu, Paras Jain, Matthew A. Wright, Azalia Mirhoseini et al.NeurIPS 2021 · 450 citations
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 335 citations
Builds on4
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement LearningHanrui Wang, Kuan Wang, Jiacheng Yang, Linxiao Shen et al.DAC 2020 · 326 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- APQ: Joint Search for Network Architecture, Pruning and Quantization PolicyTianzhe Wang, Kuan Wang, Han Cai, Ji Lin et al.CVPR 2020
Related papers
- SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 InferenceXudong Wang, Li Lyna Zhang, Jiahang Xu, Quanlu Zhang et al.ICCV 2023 · 3 citations
- LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language ModelsMojan Javaheripi, Gustavo de Rosa, Subhabrata Mukherjee, Shital Shah et al.NeurIPS 2022 · 27 citations
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 32 citations
- Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designHongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee et al.MICRO 2022 · 63 citations
- Vector-Vector-Matrix Architecture: A Novel Hardware-Aware Framework for Low-Latency Inference in NLP ApplicationsMatthew Khoury, Rumen Dangovski, Longwu Ou, Preslav Nakov et al.EMNLP 2020 · 2 citations
