Few-Shot Object Detection with Fully Cross-Transformer
Guangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen, Shih-Fu Chang
Abstract
Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and calculate the similarity between image regions and few-shot examples for detection. However, in previous works, the interaction between the two branches is only restricted in the detection head, while leaving the remaining hundreds of layers for separate feature extraction. Inspired by the recent work on vision transformers and vision-language transformers, we propose a novel Fully Cross-Transformer based model (FCT) for FSOD by incorporating cross-transformer into both the feature backbone and detection head. The asymmetric-batched cross-attention is proposed to aggregate the key information from the two branches with different batch sizes. Our model can improve the few-shot similarity learning between the two branches by introducing the multi-level interactions. Comprehensive experiments on both PASCAL VOC and MSCOCO FSOD benchmarks demonstrate the effectiveness of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e78ab15e-1589-4769-9e12-42ff09adfcd9Cited by top-tier papers31
- Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature AlignmentGuangxing Han, Shiyuan Huang, Jiawei Ma, Yicheng He et al.AAAI 2022 · 227 citations
- Multi-modal Queried Object Detection in the WildYifan Xu, Mengdan Zhang, Chaoyou Fu, Peixian Chen et al.NeurIPS 2023 · 73 citations
- Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object DetectionXiaonan Lu, Wenhui Diao, Yongqiang Mao, Junxi Li et al.AAAI 2023 · 66 citations
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 61 citations
- Fine-Grained Prototypes Distillation for Few-Shot Object DetectionZichen Wang, Bo Yang, Haonan Yue, Zhenghao MaAAAI 2024 · 55 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Few-Shot Object Detection with Foundation ModelsGuangxing Han, Ser-Nam LimCVPR 2024
- Query Adaptive Few-Shot Object Detection with Heterogeneous Graph Convolutional NetworksGuangxing Han, Yicheng He, Shiyuan Huang, Jiawei Ma et al.ICCV 2021 · 134 citations
- Dense Relation Distillation With Context-Aware Aggregation for Few-Shot Object DetectionHanzhe Hu, Shuai Bai, Aoxue Li, Jinshi Cui et al.CVPR 2021
- CAD: Co-Adapting Discriminative Features for Improved Few-Shot ClassificationPhilip Chikontwe, Soopil Kim, Sang Hyun ParkCVPR 2022 · 46 citations
- Context-Transformer: Tackling Object Confusion for Few-Shot DetectionZe Yang, Yali Wang, Xianyu Chen, Jianzhuang Liu et al.AAAI 2020 · 91 citations
