Slicing Vision Transformer for Flexible Inference
Yitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang, Ke Ma, Xi Stephen Chen, Derek Hao Hu, Yun Fu
摘要
Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe that smaller ViTs are intrinsically the sub-networks of a larger ViT with different widths. Thus, we propose a general framework, named Scala, to enable a single network to represent multiple smaller ViTs with flexible inference capability, which aligns with the inherent design of ViT to vary from widths. Concretely, Scala activates several subnets during training, introduces Isolated Activation to disentangle the smallest sub-network from other subnets, and leverages Scale Coordination to ensure each sub-network receives simplified, steady, and accurate learning objectives. Comprehensive empirical validations on different tasks demonstrate that with only one-shot training, Scala learns slimmable representation without modifying the original ViT structure and matches the performance of Separate Training. Compared with the prior art, Scala achieves an average improvement of 1.6% on ImageNet-1K with fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- EA-Vit: Efficient Adaptation for Elastic Vision TransformerChen Zhu, Wangbo Zhao, Huiwen Zhang, Yuhao Zhou 等ICCV 2025 · 被引用 2 次
- HydraViT: Stacking Heads for a Scalable ViTJanek Haberer, Ali Hojjat, Olaf LandsiedelNeurIPS 2024 · 被引用 11 次
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang 等NeurIPS 2024 · 被引用 7 次
- Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization SpaceArnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu 等CVPR 2022 · 被引用 64 次
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song 等ICLR 2022 · 被引用 27 次
