Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer
Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, Xing Sun
摘要
Vision transformers (ViTs) have recently received explosive popularity, but the huge computational cost is still a severe issue. Since the computation complexity of ViT is quadratic with respect to the input sequence length, a mainstream paradigm for computation reduction is to reduce the number of tokens. Existing designs include structured spatial compression that uses a progressive shrinking pyramid to reduce the computations of large feature maps, and unstructured token pruning that dynamically drops redundant tokens. However, the limitation of existing token pruning lies in two folds: 1) the incomplete spatial structure caused by pruning is not compatible with structured spatial compression that is commonly used in modern deep-narrow transformers; 2) it usually requires a time-consuming pre-training procedure. To tackle the limitations and expand the applicable scenario of token pruning, we present Evo-ViT, a self-motivated slow-fast token evolution approach for vision transformers. Specifically, we conduct unstructured instance-wise token selection by taking advantage of the simple and effective global class attention that is native to vision transformers. Then, we propose to update the selected informative tokens and uninformative tokens with different computation paths, namely, slow-fast updating. Since slow-fast updating mechanism maintains the spatial structure and information flow, Evo-ViT can accelerate vanilla transformers of both flat and deep-narrow structures from the very beginning of the training process. Experimental results demonstrate that our method significantly reduces the computational cost of vision transformers while maintaining comparable performance on image classification. For example, our method accelerates DeiT-S by over 60% throughput while only sacrificing 0.4% top-1 accuracy on ImageNet-1K, outperforming current token pruning methods on both accuracy and efficiency. Code is available at https://github.com/YifanXu74/Evo-ViT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- StyTr2: Image Style Transfer with TransformersYingying Deng, Fan Tang, Weiming Dong, Chongyang Ma 等CVPR 2022 · 被引用 345 次
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 被引用 193 次
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao 等NeurIPS 2022 · 被引用 185 次
- Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic ChipsMan Yao, Jiakui Hu, Tianxiang Hu, Yifan Xu 等ICLR 2024 · 被引用 154 次
- FasterViT: Fast Vision Transformers with Hierarchical AttentionAli Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao 等ICLR 2024 · 被引用 132 次
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
相关 Paper
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song 等ICLR 2022 · 被引用 137 次
- Making Vision Transformers Efficient from A Token Sparsification ViewShuning Chang, Pichao Wang, Ming Lin, Fan Wang 等CVPR 2023
- ImagePiece: Content-aware Re-tokenization for Efficient Image RecognitionSeungdong Yoa, Seungjun Lee, Hye-Seung Cho, Bumsoo Kim 等AAAI 2025 · 被引用 1 次
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya 等CVPR 2022 · 被引用 288 次
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie 等HPCA 2023 · 被引用 117 次
