Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image Classification
Bo Zhang, Jiakang Yuan, Baopu Li, Tao Chen, Jiayuan Fan, Botian Shi
摘要
Few-shot fine-grained learning aims to classify a query image into one of a set of support categories with fine-grained differences. Although learning different objects' local differences via Deep Neural Networks has achieved success, how to exploit the query-support cross-image object semantic relations in Transformer-based architecture remains under-explored in the few-shot fine-grained scenario. In this work, we propose a Transformer-based doublehelix model, namely HelixFormer, to achieve the cross-image object semantic relation mining in a bidirectional and symmetrical manner. The HelixFormer consists of two steps: 1) Relation Mining Process (RMP) across different branches, and 2) Representation Enhancement Process (REP) within each individual branch. By the designed RMP, each branch can extract fine-grained object-level Cross-image Semantic Relation Maps (CSRMs) using information from the other branch, ensuring better cross-image interaction in semantically related local object regions. Further, with the aid of CSRMs, the developed REP can strengthen the extracted features for those discovered semantically-related local regions in each branch, boosting the model's ability to distinguish subtle feature differences of fine-grained objects. Extensive experiments conducted on five public fine-grained benchmarks demonstrate that HelixFormer can effectively enhance the cross-image object semantic relation matching for recognizing fine-grained objects, achieving much better performance over most state-of-the-art methods under 1-shot and 5-shot scenarios. Our code is available at: https:// github.com/ JiakangYuan/ HelixFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Cross-Layer and Cross-Sample Feature Optimization Network for Few-Shot Fine-Grained Image ClassificationZhen-Xiang Ma, Zhen-Duo Chen, Li-Jun Zhao, Zi-Chao Zhang 等AAAI 2024 · 被引用 57 次
- Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationShicheng Yang, Xiaoxu Li, Dongliang Chang, Zhanyu Ma 等ACM MM 2024 · 被引用 12 次
- Few-Shot Fine-Grained Image Classification with Progressively Feature Refinement and Continuous Relationship ModelingZhen-Xiang Ma, Zhen-Duo Chen, Tai Zheng, Xin Luo 等AAAI 2025 · 被引用 8 次
- Deciphering Perceptual Quality in Colored Point Cloud: Prioritizing Geometry or Texture Distortion?Xuemei Zhou, Irene Viola, Yunlu Chen, Jiahuan Pei 等ACM MM 2024 · 被引用 4 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
相关 Paper
- Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationYike Wu, Bo Zhang, Gang Yu, Weixi Zhang 等ACM MM 2021 · 被引用 40 次
- Focus on Query: Adversarial Mining Transformer for Few-Shot SegmentationYuan Wang, Naisong Luo, Tianzhu ZhangNeurIPS 2023 · 被引用 29 次
- SpatialFormer: Semantic and Target Aware Attentions for Few-Shot LearningJinxiang Lai, Siqian Yang, Wenlong Wu, Tao Wu 等AAAI 2023 · 被引用 21 次
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 被引用 43 次
- Enhancing Transformer-based Semantic Matching for Few-shot Learning through Weakly Contrastive Pre-trainingWei Yang, Tengfei Huo, Zhiqiang LiuACM MM 2024
