MIA-Former: Efficient and Robust Vision Transformers via Multi-Grained Input-Adaptation
Zhongzhi Yu, Yonggan Fu, Sicheng Li, Chaojian Li, Yingyan Lin
摘要
Vision transformers have recently demonstrated great success in various computer vision tasks, motivating a tremendously increased interest in their deployment into many real-world IoT applications. However, powerful ViTs are often too computationally expensive to be fitted onto real-world resource-constrained platforms, due to (1) their quadratically increased complexity with the number of input tokens and (2) their overparameterized self-attention heads and model depth. In parallel, different images are of varied complexity and their different regions can contain various levels of visual information, e.g., a sky background is not as informative as a foreground object in object classification tasks, indicating that treating those regions equally in terms of model complexity is unnecessary while such opportunities for trimming down ViTs' complexity have not been fully exploited. To this end, we propose a Multi-grained Input-Adaptive Vision Transformer framework dubbed MIA-Former that can input-adaptively adjust the structure of ViTs at three coarse-to-fine-grained granularities (i.e., model depth and the number of model heads/tokens). In particular, our MIA-Former adopts a low-cost network trained with a hybrid supervised and reinforcement learning method to skip the unnecessary layers, heads, and tokens in an input adaptive manner, reducing the overall computational cost. Furthermore, an interesting side effect of our MIA-Former is that its resulting ViTs are naturally equipped with improved robustness against adversarial attacks over their static counterparts, because MIA-Former's multi-grained dynamic control improves the model diversity similar to the effect of ensemble and thus increases the difficulty of adversarial attacks against all its sub-models. Extensive experiments and ablation studies validate that the proposed MIA-Former framework can (1) effectively allocate adaptive computation budgets to the difficulty of input images, achieving state-of-the-art (SOTA) accuracy-efficiency trade-offs, e.g., up to 16.5% computation savings with the same or even a higher accuracy compared with the SOTA dynamic transformer models, and (2) boost ViTs' robustness accuracy under various adversarial attacks over their vanilla counterparts by 2.4% and 3.0%, respectively. Our code is available at https://github.com/RICE-EIC/MIA-Former.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision TransformerHaoran You, Huihong Shi, Yipin Guo, Yingyan LinNeurIPS 2023 · 被引用 27 次
- Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsHalima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar 等DAC 2023 · 被引用 14 次
- NetBooster: Empowering Tiny Deep Learning By Standing on the Shoulders of Deep GiantsZhongzhi Yu, Yonggan Fu, Jiayi Yuan, Haoran You 等DAC 2023 · 被引用 1 次
- Hint-Aug: Drawing Hints from Foundation Vision Transformers towards Boosted Few-shot Parameter-Efficient TuningZhongzhi Yu, Shang Wu, Yonggan Fu, Shunyao Zhang 等CVPR 2023
- Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer InferenceHaoran You, Yunyang Xiong, Xiaoliang Dai, Bichen Wu 等CVPR 2023
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
相关 Paper
- AdaViT: Adaptive Vision Transformers for Efficient Image RecognitionLingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan 等CVPR 2022
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 被引用 70 次
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya 等CVPR 2022 · 被引用 288 次
- TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationWenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen 等CVPR 2022 · 被引用 313 次
- BiFormer: Vision Transformer with Bi-Level Routing AttentionLei Zhu, Xinjiang Wang, Zhanghan Ke, Wayne Zhang 等CVPR 2023
