Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training
Zhenglun Kong, Haoyu Ma, Geng Yuan, Mengshu Sun, Yanyue Xie, Peiyan Dong, Xin Meng, Xuan Shen, Hao Tang, Minghai Qin, Tianlong Chen, Xiaolong Ma
Abstract
Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit their generalization. Previous compression algorithms usually start from the pre-trained dense models and only focus on efficient inference, while time-consuming training is still unavoidable. In contrast, this paper points out that the million-scale training data is redundant, which is the fundamental reason for the tedious training. To address the issue, this paper aims to introduce sparsity into data and proposes an end-to-end efficient training framework from three sparse perspectives, dubbed Tri-Level E-ViT. Specifically, we leverage a hierarchical data redundancy reduction scheme, by exploring the sparsity under three levels: number of training examples in the dataset, number of patches (tokens) in each example, and number of connections between tokens that lie in attention weights. With extensive experiments, we demonstrate that our proposed technique can noticeably accelerate training for various ViT architectures while maintaining accuracy. Remarkably, under certain ratios, we are able to improve the ViT accuracy rather than compromising it. For example, we can achieve 15.2% speedup with 72.6% (+0.4) Top-1 accuracy on Deit-T, and 15.7% speedup with 79.9% (+0.1) Top-1 accuracy on Deit-S. This proves the existence of data redundancy in ViT. Our code is released at https://github.com/ZLKong/Tri-Level-ViT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48c21b54-a282-44a4-a5cd-351bf8c8ff52Cited by top-tier papers11
- Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the EdgeXuan Shen, Peiyan Dong, Lei Lu, Zhenglun Kong et al.AAAI 2024 · 59 citations
- DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and RoutingConglong Li, Zhewei Yao, Xiaoxia Wu, Minjia Zhang et al.AAAI 2024 · 43 citations
- LazyDiT: Lazy Learning for the Acceleration of Diffusion TransformersXuan Shen, Zhao Song, Yufa Zhou, Bo Chen et al.AAAI 2025 · 40 citations
- TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing ModalitiesZheyu Zhang, Gang Yang, Yueyi Zhang, Huanjing Yue et al.AAAI 2024 · 29 citations
- Numerical Pruning for Efficient Autoregressive ModelsXuan Shen, Zhao Song, Yufa Zhou, Bo Chen et al.AAAI 2025 · 5 citations
Builds on31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah et al.CVPR 2022 · 36 citations
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
