SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning
Jinxiang Lai, Siqian Yang, Wenlong Wu, Tao Wu, Guannan Jiang, Xi Wang, Jun Liu, Bin-Bin Gao, Wei Zhang, Yuan Xie, Chengjie Wang
Abstract
Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar regions of support and query pairs. However, it suffers from two problems: CNN structure produces inaccurate attention map based on local features, and mutually similar backgrounds cause distraction. To alleviate these problems, we design a novel SpatialFormer structure to generate more accurate attention regions based on global features. Different from the traditional Transformer modeling intrinsic instance-level similarity which causes accuracy degradation in FSL, our SpatialFormer explores the semantic-level similarity between pair inputs to boost the performance. Then we derive two specific attention modules, named SpatialFormer Semantic Attention (SFSA) and SpatialFormer Target Attention (SFTA), to enhance the target object regions while reduce the background distraction. Particularly, SFSA highlights the regions with same semantic information between pair features, and SFTA finds potential foreground object regions of novel feature that are similar to base categories. Extensive experiments show that our methods are effective and achieve new state-of-the-art results on few-shot classification benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44b7e04b-9d42-40ea-a832-e0b71443d3a1Cited by top-tier papers1
Ask how each one uses itBuilds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 420 citations
- Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot ClassificationJiangtao Xie, Fei Long, Jiaming Lv, Qilong Wang et al.CVPR 2022 · 270 citations
- Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot LearningZhiqiang Shen, Zechun Liu, Jie Qin, Marios Savvides et al.AAAI 2021 · 203 citations
Related papers
- Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationYike Wu, Bo Zhang, Gang Yu, Weixi Zhang et al.ACM MM 2021 · 40 citations
- CAD: Co-Adapting Discriminative Features for Improved Few-Shot ClassificationPhilip Chikontwe, Soopil Kim, Sang Hyun ParkCVPR 2022 · 46 citations
- Integrative Few-Shot Learning for Classification and SegmentationDahyun Kang, Minsu ChoCVPR 2022 · 76 citations
- Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationShicheng Yang, Xiaoxu Li, Dongliang Chang, Zhanyu Ma et al.ACM MM 2024 · 12 citations
- Weak-shot Semantic Segmentation via Dual Similarity TransferJunjie Chen, Li Niu, Siyuan Zhou, Jianlou Si et al.NeurIPS 2022 · 15 citations
