Self-Calibrated Cross Attention Network for Few-Shot Segmentation
Qianxiong Xu, Wenting Zhao, Guosheng Lin, Cheng Long
Abstract
The key to the success of few-shot segmentation (FSS) lies in how to effectively utilize support samples. Most solutions compress support foreground (FG) features into prototypes, but lose some spatial details. Instead, others use cross attention to fuse query features with uncompressed support FG. Query FG could be fused with support FG, however, query background (BG) cannot find matched BG features in support FG, yet inevitably integrates dissimilar features. Besides, as both query FG and BG are combined with support FG, they get entangled, thereby leading to ineffective segmentation. To cope with these issues, we design a self-calibrated cross attention (SCCA) block. For efficient patch-based attention, query and support features are firstly split into patches. Then, we design a patch alignment module to align each query patch with its most similar support patch for better cross attention. Specifically, SCCA takes a query patch as Q, and groups the patches from the same query image and the aligned patches from the support image as K&V . In this way, the query BG features are fused with matched BG features (from query patches), and thus the aforementioned issues will be mitigated. Moreover, when calculating SCCA, we design a scaled-cosine mechanism to better utilize the support features for similarity calculation. Extensive experiments conducted on PASCAL-5 i and COCO-20 i demonstrate the superiority of our model, e.g., the mIoU score under 5-shot setting on COCO-20 i is 5.6%+ better than previous stateof-the-arts. The code is available at https://github. com/Sam1224/SCCAN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80e27ba3-e120-42e5-afe5-053d31841684Cited by top-tier papers18
- Bridge the Points: Graph-based Few-shot Segment Anything SemanticallyAnqi Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu et al.NeurIPS 2024 · 56 citations
- Hybrid Mamba for Few-Shot SegmentationQianxiong Xu, Xuanyi Liu, Lanyun Zhu, Guosheng Lin et al.NeurIPS 2024 · 49 citations
- LLaFS: When Large Language Models Meet Few-Shot SegmentationLanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye et al.CVPR 2024 · 39 citations
- Cross-modulated Attention Transformer for RGBT TrackingYun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li et al.AAAI 2025 · 28 citations
- Addressing Background Context Bias in Few-Shot Segmentation Through Iterative ModulationLanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See et al.CVPR 2024 · 20 citations
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot SegmentationSuho Park, SuBeen Lee, Hyun Seok Seong, Jaejoon Yoo et al.AAAI 2025 · 9 citations
- Unlocking the Power of SAM 2 for Few-Shot SegmentationQianxiong Xu, Lanyun Zhu, Xuanyi Liu, Guosheng Lin et al.ICML 2025
- Self-Guided and Cross-Guided Learning for Few-Shot SegmentationBingfeng Zhang, Jimin Xiao, Terry QinCVPR 2021
- Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot SegmentationTianyu Zou, Shengwu Xiong, Ruilin Yao, Yi RongICCV 2025 · 5 citations
- Scale-Aware Graph Neural Network for Few-Shot Semantic SegmentationGuo-Sen Xie, Jie Liu, Huan Xiong, Ling ShaoCVPR 2021
