DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image Segmentation
Zhehao Wang, Xian Lin, Nannan Wu, Li Yu, Kwang-Ting Cheng, Zengqiang Yan
Abstract
Despite the great potential in capturing long-range dependency, one rarely-explored underlying issue of transformer in medical image segmentation is attention collapse, making it often degenerate into a bypass module in CNN-Transformer hybrid architectures. This is due to the high computational complexity of vision transformers requiring extensive training data while well-annotated medical image data is relatively limited, resulting in poor convergence. In this paper, we propose a plug-n-play transformer block with dynamic token merging, named DTMFormer, to avoid building long-range dependency on redundant and duplicated tokens and thus pursue better convergence. Specifically, DTMFormer consists of an attention-guided token merging (ATM) module to adaptively cluster tokens into fewer semantic tokens based on feature and dependency similarity and a light token reconstruction module to fuse ordinary and semantic tokens. In this way, as self-attention in ATM is calculated based on fewer tokens, DTMFormer is of lower complexity and more friendly to converge. Extensive experiments on publicly-available datasets demonstrate the effectiveness of DTMFormer working as a plug-n-play module for simultaneous complexity reduction and performance improvement. We believe it will inspire future work on rethinking transformers in medical image segmentation. Code: https://github.com/iam-nacl/DTMFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AIF-SFDA: Autonomous Information Filter Driven Source-Free Domain Adaptation for Medical Image SegmentationHaojin Li, Heng Li, Jianyu Chen, Rihan Zhong et al.AAAI 2025 · 5 citations
- Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters MoreZichen Wen, Yifeng Gao, Shaobo Wang, Junyuan Zhang et al.EMNLP 2025 · 4 citations
- Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level TasksCheng Lei, Ao Li, Hu Yao, Ce Zhu et al.CVPR 2025
- Saliency-Driven Token Merging for Vision TransformersWeiying Xie, Xiaoyu Chen, Xin Zhang, Chenhe Hao et al.CVPR 2026
- Similarity Memory Prior is All You Need for Medical Image SegmentationHao Tang, Zhiqing Guo, Liejun Wang, Chao LiuICCV 2025
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya et al.CVPR 2022 · 288 citations
- TokenLearner: Adaptive Space-Time Tokenization for VideosMichael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani et al.NeurIPS 2021 · 274 citations
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy et al.ICML 2020 · 260 citations
Related papers
- ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image SegmentationHuimin Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong et al.AAAI 2023 · 5 citations
- Class-Aware Adversarial Transformers for Medical Image SegmentationChenyu You, Ruihan Zhao, Fenglin Liu, Siyuan Dong et al.NeurIPS 2022 · 137 citations
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang et al.ICCV 2023 · 36 citations
- D3ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMsShuochen Chang, Xiaofeng Zhang, Qingyang Liu, Li NiuAAAI 2026
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationHaipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang et al.CVPR 2025
