Lune

ACM MM2025顶会

TF-ATM: Training-Free Adaptive Token Merging

Xin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen, Tianlin Hui, Jitao Ma, Leyuan Fang

2025年份
2顶会引用

摘要

Vision Transformers (ViTs) show promising potential in various multimedia application scenarios. To facilitate their deployment on resource-constrained devices, token pruning and merging have been introduced. However, existing token compression methods focus solely on the abstract features exhibited by high-dimensional tokens after patch embedding, resulting in information loss when evaluating token importance. In this paper, we propose a novel Training-Free Adaptive Token Merging (TF-ATM) method by exploring the intrinsic properties of images themselves. Our TF-ATM is inspired by the observation that the characteristics of patches can intuitively reflect the redundancy level of images. Based on the observation, we develop a method that is mathematically formulated to merge tokens corresponding to patches that are close to the Median presentation in the Frequency domain (MF). The principle behind our merging is that patches close to MF can be replaced by their surrounding ones, and thus removing them does not impair performance. Besides, we experimentally show that patches farther from MF contain more important information, which can be leveraged to capture the object of interest accurately. Without any retraining, TF-ATM leads to significant improvements over the state-of-the-arts (SOTAs), with similar FLOPs (floating point operations). For example, we achieve a 44.5%-FLOPs reduction with only a small loss of 0.39% in top-1 accuracy for the MAE-H model on ImageNet dataset, superior to comparison approaches that require meticulous fine-tuning.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 49bde12f-7a2f-4c46-bdb2-7bf79c160e7f

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖