TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen, Tianlin Hui, Jitao Ma, Leyuan Fang
Abstract
Vision Transformers (ViTs) show promising potential in various multimedia application scenarios. To facilitate their deployment on resource-constrained devices, token pruning and merging have been introduced. However, existing token compression methods focus solely on the abstract features exhibited by high-dimensional tokens after patch embedding, resulting in information loss when evaluating token importance. In this paper, we propose a novel Training-Free Adaptive Token Merging (TF-ATM) method by exploring the intrinsic properties of images themselves. Our TF-ATM is inspired by the observation that the characteristics of patches can intuitively reflect the redundancy level of images. Based on the observation, we develop a method that is mathematically formulated to merge tokens corresponding to patches that are close to the Median presentation in the Frequency domain (MF). The principle behind our merging is that patches close to MF can be replaced by their surrounding ones, and thus removing them does not impair performance. Besides, we experimentally show that patches farther from MF contain more important information, which can be leveraged to capture the object of interest accurately. Without any retraining, TF-ATM leads to significant improvements over the state-of-the-arts (SOTAs), with similar FLOPs (floating point operations). For example, we achieve a 44.5%-FLOPs reduction with only a small loss of 0.39% in top-1 accuracy for the MAE-H model on ImageNet dataset, superior to comparison approaches that require meticulous fine-tuning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 49bde12f-7a2f-4c46-bdb2-7bf79c160e7fCited by top-tier papers2
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong et al.ICML 2026
- Saliency-Driven Token Merging for Vision TransformersWeiying Xie, Xiaoyu Chen, Xin Zhang, Chenhe Hao et al.CVPR 2026
Related papers
- Spatially-Regularized Entropy for Discriminative Token Merging in Fine-Grained Re-IdentificationShangze Li, Yifan Xu, Jingmiao Liang, Yongfei Zhang et al.ICML 2026
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersSifan Long, Zhen Zhao, Jimin Pi, Shengsheng Wang et al.CVPR 2023
- vid-TLDR: Training Free Token merging for Light-Weight Video TransformerJoonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi et al.CVPR 2024
- EViT: Expediting Vision Transformers via Token ReorganizationsYouwei Liang, Chongjian Ge, Zhan Tong, Yibing Song et al.ICLR 2022 · 137 citations
