Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image Classification
Bo Zhang, Jiakang Yuan, Baopu Li, Tao Chen, Jiayuan Fan, Botian Shi
Abstract
Few-shot fine-grained learning aims to classify a query image into one of a set of support categories with fine-grained differences. Although learning different objects' local differences via Deep Neural Networks has achieved success, how to exploit the query-support cross-image object semantic relations in Transformer-based architecture remains under-explored in the few-shot fine-grained scenario. In this work, we propose a Transformer-based doublehelix model, namely HelixFormer, to achieve the cross-image object semantic relation mining in a bidirectional and symmetrical manner. The HelixFormer consists of two steps: 1) Relation Mining Process (RMP) across different branches, and 2) Representation Enhancement Process (REP) within each individual branch. By the designed RMP, each branch can extract fine-grained object-level Cross-image Semantic Relation Maps (CSRMs) using information from the other branch, ensuring better cross-image interaction in semantically related local object regions. Further, with the aid of CSRMs, the developed REP can strengthen the extracted features for those discovered semantically-related local regions in each branch, boosting the model's ability to distinguish subtle feature differences of fine-grained objects. Extensive experiments conducted on five public fine-grained benchmarks demonstrate that HelixFormer can effectively enhance the cross-image object semantic relation matching for recognizing fine-grained objects, achieving much better performance over most state-of-the-art methods under 1-shot and 5-shot scenarios. Our code is available at: https:// github.com/ JiakangYuan/ HelixFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e8797bd-18d9-46cb-887a-a67d8273210cCited by top-tier papers4
- Cross-Layer and Cross-Sample Feature Optimization Network for Few-Shot Fine-Grained Image ClassificationZhen-Xiang Ma, Zhen-Duo Chen, Li-Jun Zhao, Zi-Chao Zhang et al.AAAI 2024 · 57 citations
- Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationShicheng Yang, Xiaoxu Li, Dongliang Chang, Zhanyu Ma et al.ACM MM 2024 · 12 citations
- Few-Shot Fine-Grained Image Classification with Progressively Feature Refinement and Continuous Relationship ModelingZhen-Xiang Ma, Zhen-Duo Chen, Tai Zheng, Xin Luo et al.AAAI 2025 · 8 citations
- Deciphering Perceptual Quality in Colored Point Cloud: Prioritizing Geometry or Texture Distortion?Xuemei Zhou, Irene Viola, Yunlu Chen, Jiahuan Pei et al.ACM MM 2024 · 4 citations
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
Related papers
- Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationYike Wu, Bo Zhang, Gang Yu, Weixi Zhang et al.ACM MM 2021 · 40 citations
- Focus on Query: Adversarial Mining Transformer for Few-Shot SegmentationYuan Wang, Naisong Luo, Tianzhu ZhangNeurIPS 2023 · 29 citations
- SpatialFormer: Semantic and Target Aware Attentions for Few-Shot LearningJinxiang Lai, Siqian Yang, Wenlong Wu, Tao Wu et al.AAAI 2023 · 21 citations
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 43 citations
- Enhancing Transformer-based Semantic Matching for Few-shot Learning through Weakly Contrastive Pre-trainingWei Yang, Tengfei Huo, Zhiqiang LiuACM MM 2024
