Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization Space
Arnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu, Kwang-Ting Cheng, Eric P. Xing
摘要
This paper explores the feasibility of finding an optimal sub-model from a vision transformer and introduces a pure vision transformer slimming (ViT-Slim) framework. It can search a sub-structure from the original model end-to-end across multiple dimensions, including the input tokens, MHSA and MLP modules with state-of-the-art performance. Our method is based on a learnable and unified ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> sparsity constraint with pre-defined factors to reflect the global importance in the continuous searching space of different dimensions. The searching process is highly efficient through a single-shot training scheme. For instance, on DeiT-S, ViT-Slim only takes 43 GPU hours for the searching process, and the searched structure is flexible with diverse dimensionalities in different modules. Then, a budget threshold is employed according to the requirements of accuracy-FLOPs trade-off on running devices, and a retraining process is performed to obtain the final model. The extensive experiments show that our ViT-Slim can compress up to 40% of parameters and 40% FLOPs on various vision transformers while increasing the accuracy by 0.6% on ImageNet. We also demonstrate the advantage of our searched models on several downstream datasets. Our code is available at https://github.com/Arnav0400/ViT-Slim.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis 等ICCV 2023 · 被引用 300 次
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 被引用 193 次
- TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight InheritanceKan Wu, Houwen Peng, Zhenghong Zhou, Bin Xiao 等ICCV 2023 · 被引用 118 次
- SAViT: Structure-Aware Vision Transformer Pruning via Collaborative OptimizationChuanyang Zheng, Zheyang Li, Kai Zhang, Zhi Yang 等NeurIPS 2022 · 被引用 98 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan 等NeurIPS 2021 · 被引用 295 次
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan 等ICLR 2022 · 被引用 118 次
- Patch Slimming for Efficient Vision TransformersYehui Tang, Kai Han, Yunhe Wang, Chang Xu 等CVPR 2022 · 被引用 173 次
- Slicing Vision Transformer for Flexible InferenceYitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang 等NeurIPS 2024 · 被引用 4 次
- Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision TransformersSiyuan Wei, Tianzhu Ye, Shen Zhang, Yao Tang 等CVPR 2023
