Adder Attention for Vision Transformer
Han Shu, Jiahao Wang, Hanting Chen, Lin Li, Yujiu Yang, Yunhe Wang
Abstract
Transformer is a new kind of calculation paradigm for deep learning which has shown strong performance on a large variety of computer vision tasks. However, compared with conventional deep models (e.g., convolutional neural networks), vision transformers require more computational resources which cannot be easily deployed on mobile devices. To this end, we present to reduce the energy consumptions using adder neural network (AdderNet). We first theoretically analyze the mechanism of self-attention and the difficulty for applying adder operation into this module. Specifically, the feature diversity, i.e., the rank of attention map using only additions cannot be well preserved. Thus, we develop an adder attention layer that includes an additional identity mapping. With the new operation, vision transformers constructed using additions can also provide powerful feature representations. Experimental results on several benchmarks demonstrate that the proposed approach can achieve highly competitive performance to that of the baselines while achieving an about 2˜3× reduction on the energy consumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ff2db0c-943a-494e-bec4-40a93fd957fbCited by top-tier papers9
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou et al.NeurIPS 2024 · 47 citations
- EcoFormer: Energy-Saving Attention with Linear ComplexityJing Liu, Zizheng Pan, Haoyu He, Jianfei Cai et al.NeurIPS 2022 · 38 citations
- ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision TransformerHaoran You, Huihong Shi, Yipin Guo, Yingyan LinNeurIPS 2023 · 27 citations
- Multiplication-Free Transformer Training via Piecewise Affine OperationsAtli Kosson, Martin JaggiNeurIPS 2023 · 16 citations
- Redistribution of Weights and Activations for AdderNet QuantizationYing Nie, Kai Han, Haikang Diao, Chuanjian Liu et al.NeurIPS 2022 · 14 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- AdderSR: Towards Energy Efficient Image Super-ResolutionDehua Song, Yunhe Wang, Hanting Chen, Chang Xu et al.CVPR 2021
- Winograd Algorithm for AdderNetWenshuo Li, Hanting Chen, Mingqiang Huang, Xinghao Chen et al.ICML 2021 · 8 citations
- An Empirical Study of Adder Neural Networks for Object DetectionXinghao Chen, Chang Xu, Minjing Dong, Chunjing Xu et al.NeurIPS 2021 · 22 citations
- Exploring Salient Object Detection with Adder Neural NetworksBo-Wen Yin, Zheng LinAAAI 2025 · 4 citations
- Towards Stable and Robust AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 11 citations
