Slicing Vision Transformer for Flexible Inference
Yitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang, Ke Ma, Xi Stephen Chen, Derek Hao Hu, Yun Fu
Abstract
Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe that smaller ViTs are intrinsically the sub-networks of a larger ViT with different widths. Thus, we propose a general framework, named Scala, to enable a single network to represent multiple smaller ViTs with flexible inference capability, which aligns with the inherent design of ViT to vary from widths. Concretely, Scala activates several subnets during training, introduces Isolated Activation to disentangle the smallest sub-network from other subnets, and leverages Scale Coordination to ensure each sub-network receives simplified, steady, and accurate learning objectives. Comprehensive empirical validations on different tasks demonstrate that with only one-shot training, Scala learns slimmable representation without modifying the original ViT structure and matches the performance of Separate Training. Compared with the prior art, Scala achieves an average improvement of 1.6% on ImageNet-1K with fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5e27fa3-3e4c-4331-a683-0d5f4245fa01Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- EA-Vit: Efficient Adaptation for Elastic Vision TransformerChen Zhu, Wangbo Zhao, Huiwen Zhang, Yuhao Zhou et al.ICCV 2025 · 2 citations
- HydraViT: Stacking Heads for a Scalable ViTJanek Haberer, Ali Hojjat, Olaf LandsiedelNeurIPS 2024 · 11 citations
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang et al.NeurIPS 2024 · 7 citations
- Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization SpaceArnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu et al.CVPR 2022 · 64 citations
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song et al.ICLR 2022 · 27 citations
