A General and Efficient Training for Transformer via Token Expansion
Wenxuan Huang, Yunhang Shen, Jiao Xie, Baochang Zhang, Gaoqi He, Ke Li, Xing Sun, Shaohui Lin
Abstract
The remarkable performance of Vision Transformers (ViTs) typically requires an extremely large training cost. Existing methods have attempted to accelerate the training of ViTs, yet typically disregard method universality with accuracy dropping. Meanwhile, they break the training consistency of the original transformers, including the consistency of hyper-parameters, architecture, and strategy, which prevents them from being widely applied to different Transformer networks. In this paper, we propose a novel token growth scheme Token Expansion (termed ToE) to achieve consistent training acceleration for ViTs. We introduce an "initialization-expansion-merging" pipeline to maintain the integrity of the intermediate feature distribution of original transformers, preventing the loss of crucial learnable information in the training process. ToE can not only be seamlessly integrated into the training and finetuning process of transformers (e.g., DeiT and LV-ViT), but also effective for efficient training frameworks (e.g., Effi-cientTrain), without twisting the original training hyperparameters, architecture, and introducing additional training strategies. Extensive experiments demonstrate that ToE achieves about 1.3× faster for the training of ViTs in a lossless manner, or even with performance gains over the full-token training baselines. Code is available at https: //github.com/Osilly/TokenExpansion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20153c85-dc7c-4654-816f-ed71a172ce2fCited by top-tier papers2
- Weakly Supervised Semantic Segmentation via Progressive Confidence Region ExpansionXiangfeng Xu, Pinyi Zhang, Wenxuan Huang, Yunhang Shen et al.CVPR 2025
- Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context SparsificationWenxuan Huang, Zijie Zhai, Yunhang Shen, Shaosheng Cao et al.ICLR 2025
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
Related papers
- Automated Progressive Learning for Efficient Training of Vision TransformersChanglin Li, Bohan Zhuang, Guangrun Wang, Xiaodan Liang et al.CVPR 2022 · 28 citations
- Token Merging: Your ViT But FasterDaniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang et al.ICLR 2023 · 62 citations
- Network Expansion For Practical Training AccelerationNing Ding, Yehui Tang, Kai Han, Chao Xu et al.CVPR 2023
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song et al.ICLR 2022 · 27 citations
