Adder Attention for Vision Transformer
Han Shu, Jiahao Wang, Hanting Chen, Lin Li, Yujiu Yang, Yunhe Wang
摘要
Transformer is a new kind of calculation paradigm for deep learning which has shown strong performance on a large variety of computer vision tasks. However, compared with conventional deep models (e.g., convolutional neural networks), vision transformers require more computational resources which cannot be easily deployed on mobile devices. To this end, we present to reduce the energy consumptions using adder neural network (AdderNet). We first theoretically analyze the mechanism of self-attention and the difficulty for applying adder operation into this module. Specifically, the feature diversity, i.e., the rank of attention map using only additions cannot be well preserved. Thus, we develop an adder attention layer that includes an additional identity mapping. With the new operation, vision transformers constructed using additions can also provide powerful feature representations. Experimental results on several benchmarks demonstrate that the proposed approach can achieve highly competitive performance to that of the baselines while achieving an about 2˜3× reduction on the energy consumption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou 等NeurIPS 2024 · 被引用 47 次
- EcoFormer: Energy-Saving Attention with Linear ComplexityJing Liu, Zizheng Pan, Haoyu He, Jianfei Cai 等NeurIPS 2022 · 被引用 38 次
- ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision TransformerHaoran You, Huihong Shi, Yipin Guo, Yingyan LinNeurIPS 2023 · 被引用 27 次
- Multiplication-Free Transformer Training via Piecewise Affine OperationsAtli Kosson, Martin JaggiNeurIPS 2023 · 被引用 16 次
- Redistribution of Weights and Activations for AdderNet QuantizationYing Nie, Kai Han, Haikang Diao, Chuanjian Liu 等NeurIPS 2022 · 被引用 14 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- AdderSR: Towards Energy Efficient Image Super-ResolutionDehua Song, Yunhe Wang, Hanting Chen, Chang Xu 等CVPR 2021
- Winograd Algorithm for AdderNetWenshuo Li, Hanting Chen, Mingqiang Huang, Xinghao Chen 等ICML 2021 · 被引用 8 次
- An Empirical Study of Adder Neural Networks for Object DetectionXinghao Chen, Chang Xu, Minjing Dong, Chunjing Xu 等NeurIPS 2021 · 被引用 22 次
- Exploring Salient Object Detection with Adder Neural NetworksBo-Wen Yin, Zheng LinAAAI 2025 · 被引用 4 次
- Towards Stable and Robust AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 被引用 11 次
