Suppressing the Heterogeneity: A Strong Feature Extractor for Few-shot Segmentation
Zhengdong Hu, Yifan Sun, Yi Yang
Abstract
This paper tackles the Few-shot Semantic Segmentation (FSS) task with focus on learning the feature extractor. Somehow the feature extractor has been overlooked by recent state-of-the-art methods, which directly use a deep model pretrained on ImageNet for feature extraction (without further fine-tuning). Under this background, we think the FSS feature extractor deserves exploration and observe the heterogeneity (i.e., the intra-class diversity in the raw images) as a critical challenge hindering the intra-class feature compactness. The heterogeneity has three levels from coarse to fine: 1) Sample-level: the inevitable distribution gap between the support and query images makes them heterogeneous from each other. 2) Region-level: the background in FSS actually contains multiple regions with different semantics. 3) Patch-level: some neighboring patches belonging to a same class may appear quite different from each other. Motivated by these observations, we propose a feature extractor with Multi-level Heterogeneity Suppressing (MuHS). MuHS leverages the attention mechanism in transformer backbone to effectively suppress all these three-level heterogeneity. Concretely, MuHS reinforces the attention / interaction between different samples (query and support), different regions and neighboring patches by constructing cross-sample attention, cross-region interaction and a novel masked image segmentation (inspired by the recent masked image modeling), respectively. We empirically show that 1) MuHS brings consistent improvement for various FSS heads and 2) using a simple linear classification head, MuHS sets new states of the art on multiple FSS datasets, validating the importance of FSS feature learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 95da47d5-e075-47a5-b257-d2bff44142c7Cited by top-tier papers10
- Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image RetrievalYucheng Suo, Fan Ma, Linchao Zhu, Yi YangCVPR 2024 · 20 citations
- Learning Solution-Aware Transformers for Efficiently Solving Quadratic Assignment ProblemZhentao Tan, Yadong MuICML 2024 · 5 citations
- Object-Level Correlation for Few-Shot SegmentationChunlin Wen, Yu Zhang, Jie Fan, Hongyuan Zhu et al.ICCV 2025 · 5 citations
- Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge TransferXinyue Chen, Miaojing Shi, Zijian Zhou, Lianghua He et al.AAAI 2025 · 3 citations
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
Related papers
- Feature-Proxy Transformer for Few-Shot SegmentationJian-Wei Zhang, Yifan Sun, Yi Yang, Wei ChenNeurIPS 2022 · 105 citations
- Hierarchical Dense Correlation Distillation for Few-Shot SegmentationBohao Peng, Zhuotao Tian, Xiaoyang Wu, Chengyao Wang et al.CVPR 2023
- Simpler is Better: Few-shot Semantic Segmentation with Classifier Weight TransformerZhihe Lu, Sen He, Xiatian Zhu, Li Zhang et al.ICCV 2021 · 232 citations
- Integrative Few-Shot Learning for Classification and SegmentationDahyun Kang, Minsu ChoCVPR 2022 · 76 citations
- Semantic Prompt for Few-Shot Image RecognitionCVPR 2023
