Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning
Weicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo, Weihong Lin, Ding Jia, Zheng Zhang, Chao Zhang, Han Hu
摘要
Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in large-scale vision transformers, especially for image classification tasks. Typically, they select a small group of essential tokens according to their relevance with the [class] token, then fine-tune the weights of the vision transformer. Such fine-tuning is less practical for dense prediction due to the much heavier computation and GPU memory cost than image classification. In this paper, we focus on a more challenging problem, i.e., accelerating large-scale vision transformers for dense prediction without any additional re-training or finetuning. In response to the fact that high-resolution representations are necessary for dense prediction, we present two non-parametric operators, a token clustering layer to decrease the number of tokens and a token reconstruction layer to increase the number of tokens. The following steps are performed to achieve this: (i) we use the token clustering layer to cluster the neighboring tokens together, resulting in low-resolution representations that maintain the spatial structures; (ii) we apply the following transformer layers only to these low-resolution representations or clustered tokens; and (iii) we use the token reconstruction layer to re-create the high-resolution representations from the refined low-resolution representations. The results obtained by our method are promising on five dense prediction tasks, including object detection, semantic segmentation, panoptic segmentation, instance segmentation, and depth estimation. Accordingly, our method accelerates 40% ↑ FPS and saves 30% ↓ GFLOPs of "Segmenter+ViT-L/16" while maintaining 99.5% of the performance on ADE20K without fine-tuning the official weights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything ModelZihan Zhong, Zhiqiang Tang, Tong He, Haoyang Fang 等ICLR 2024 · 被引用 91 次
- Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationQuan Tang, Bowen Zhang, Jiajun Liu, Fagui Liu 等ICCV 2023 · 被引用 74 次
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng 等NeurIPS 2023 · 被引用 63 次
- CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language TransformersDachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang 等ICML 2024 · 被引用 46 次
- AiluRus: A Scalable ViT Framework for Dense PredictionJin Li, Yaoming Wang, Xiaopeng Zhang, Bowen Shi 等NeurIPS 2023 · 被引用 18 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Making Vision Transformers Efficient from A Token Sparsification ViewShuning Chang, Pichao Wang, Ming Lin, Fan Wang 等CVPR 2023
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- GPViT: A High Resolution Non-Hierarchical Vision Transformer with Group PropagationChenhongyi Yang, Jiarui Xu, Shalini De Mello, Elliot J. Crowley 等ICLR 2023 · 被引用 7 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
