Self-Calibrated Cross Attention Network for Few-Shot Segmentation
Qianxiong Xu, Wenting Zhao, Guosheng Lin, Cheng Long
摘要
The key to the success of few-shot segmentation (FSS) lies in how to effectively utilize support samples. Most solutions compress support foreground (FG) features into prototypes, but lose some spatial details. Instead, others use cross attention to fuse query features with uncompressed support FG. Query FG could be fused with support FG, however, query background (BG) cannot find matched BG features in support FG, yet inevitably integrates dissimilar features. Besides, as both query FG and BG are combined with support FG, they get entangled, thereby leading to ineffective segmentation. To cope with these issues, we design a self-calibrated cross attention (SCCA) block. For efficient patch-based attention, query and support features are firstly split into patches. Then, we design a patch alignment module to align each query patch with its most similar support patch for better cross attention. Specifically, SCCA takes a query patch as Q, and groups the patches from the same query image and the aligned patches from the support image as K&V . In this way, the query BG features are fused with matched BG features (from query patches), and thus the aforementioned issues will be mitigated. Moreover, when calculating SCCA, we design a scaled-cosine mechanism to better utilize the support features for similarity calculation. Extensive experiments conducted on PASCAL-5 i and COCO-20 i demonstrate the superiority of our model, e.g., the mIoU score under 5-shot setting on COCO-20 i is 5.6%+ better than previous stateof-the-arts. The code is available at https://github. com/Sam1224/SCCAN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Bridge the Points: Graph-based Few-shot Segment Anything SemanticallyAnqi Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu 等NeurIPS 2024 · 被引用 56 次
- Hybrid Mamba for Few-Shot SegmentationQianxiong Xu, Xuanyi Liu, Lanyun Zhu, Guosheng Lin 等NeurIPS 2024 · 被引用 49 次
- LLaFS: When Large Language Models Meet Few-Shot SegmentationLanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye 等CVPR 2024 · 被引用 39 次
- Cross-modulated Attention Transformer for RGBT TrackingYun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li 等AAAI 2025 · 被引用 28 次
- Addressing Background Context Bias in Few-Shot Segmentation Through Iterative ModulationLanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See 等CVPR 2024 · 被引用 20 次
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
相关 Paper
- Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot SegmentationSuho Park, SuBeen Lee, Hyun Seok Seong, Jaejoon Yoo 等AAAI 2025 · 被引用 9 次
- Unlocking the Power of SAM 2 for Few-Shot SegmentationQianxiong Xu, Lanyun Zhu, Xuanyi Liu, Guosheng Lin 等ICML 2025
- Self-Guided and Cross-Guided Learning for Few-Shot SegmentationBingfeng Zhang, Jimin Xiao, Terry QinCVPR 2021
- Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot SegmentationTianyu Zou, Shengwu Xiong, Ruilin Yao, Yi RongICCV 2025 · 被引用 5 次
- Scale-Aware Graph Neural Network for Few-Shot Semantic SegmentationGuo-Sen Xie, Jie Liu, Huan Xiong, Ling ShaoCVPR 2021
