A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction
Dongyang Liu, Meina Kan, Shiguang Shan, Xilin Chen
Abstract
Vision Transformers (ViTs) are now flourishing in the computer vision area. Despite the remarkable success, ViTs suffer from high computational costs, which greatly hinder their practical usage. Token reduction, which identifies and discards unimportant tokens during forward propagation, has then been proposed to make ViTs more efficient. For token reduction methodologies, a scoring metric is essential to distinguish between important and unimportant tokens. The attention score from the [CLS] token, which takes the responsibility to aggregate useful information and form the final output, has been established by prior works as an advantageous choice. Nevertheless, whereas the task pressure is applied at the end of the whole model, token reduction generally starts from very early blocks. Given the long distance in between, in the early blocks, [CLS] token lacks the impetus to gather task-relevant information, causing somewhat arbitrary attention allocation. This phenomenon, in turn, degrades the reliability of token scoring and substantially compromises the effectiveness of token reduction. Inspired by advances in the domain of dynamic neural networks, in this paper, we introduce Multi-Exit Token Reduction (METR), a simple romance between multi-exit architecture and token reduction-two areas previously considered orthogonal. By injecting early task pressure via multi-exit loss, the [CLS] token is spurred to collect task-related information in even early blocks, thus bolstering the credibility of [CLS] attention as a token-scoring metric. Additionally, we employ self-distillation to further refine the quality of early supervision. Extensive experiments substantiate both the existence and effectiveness of the newfound chemistry. Comparative assessments also indicate that METR outperforms state-of-the-art token reduction methods on standard benchmarks, especially under aggressive reduction ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5bff668-8a3a-4b0e-8ed2-7453e6498b30Cited by top-tier papers3
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D ParallelismYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding et al.ICML 2024 · 73 citations
- Learning to Merge Tokens via Decoupled Embedding for Efficient Vision TransformersDong Hoon Lee, Seunghoon HongNeurIPS 2024 · 26 citations
- Faster Parameter-Efficient Tuning with Token Redundancy ReductionKwonyoung Kim, Jungin Park, Jin Kim, Hyeongjun Kwon et al.CVPR 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- Multi-Criteria Token Fusion with One-Step-Ahead Attention for Efficient Vision TransformersSanghyeok Lee, Joonmyung Choi, Hyunwoo J. KimCVPR 2024
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationQuan Tang, Bowen Zhang, Jiajun Liu, Fagui Liu et al.ICCV 2023 · 74 citations
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
- You Only Need Less Attention at Each Stage in Vision TransformersShuoxi Zhang, Hanpeng Liu, Stephen Lin, Kun HeCVPR 2024 · 19 citations
