ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation
Huimin Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong, Yen-Wei Chen, Hong Wang, Yuexiang Li, Yawen Huang, Yefeng Zheng
Abstract
Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a classlevel perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and interclass interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- ACFNet: Attentional Class Feature Network for Semantic SegmentationFan Zhang, Yanqin Chen, Zhihang Li, Zhibin Hong et al.ICCV 2019 · 297 citations
- Mining Contextual Information Beyond Image for Semantic SegmentationZhenchao Jin, Tao Gong, Dongdong Yu, Qi Chu et al.ICCV 2021 · 95 citations
Related papers
- Class-Aware Adversarial Transformers for Medical Image SegmentationChenyu You, Ruihan Zhao, Fenglin Liu, Siyuan Dong et al.NeurIPS 2022 · 137 citations
- HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Bernt Schiele et al.CVPR 2023
- DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image SegmentationZhehao Wang, Xian Lin, Nannan Wu, Li Yu et al.AAAI 2024 · 14 citations
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.CVPR 2022 · 275 citations
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang et al.ICCV 2023 · 36 citations
