Semantic-aligned Fusion Transformer for One-shot Object Detection
Yizhou Zhao, Xun Guo, Yan Lu
Abstract
One-shot object detection aims at detecting novel objects according to merely one given instance. With extreme data scarcity, current approaches explore various feature fusions to obtain directly transferable meta-knowledge. Yet, their performances are often unsatisfactory. In this paper, we attribute this to inappropriate correlation methods that misalign query-support semantics by overlooking spatial structures and scale variances. Upon analysis, we leverage the attention mechanism and propose a simple but effective architecture named Semantic-aligned Fusion Transformer (SaFT) to resolve these issues. Specifically, we equip SaFT with a vertical fusion module (VFM) for cross-scale semantic enhancement and a horizontal fusion module (HFM) for cross-sample feature fusion. Together, they broaden the vision for each feature point from the support to a whole augmented feature pyramid from the query, facilitating semantic-aligned associations. Extensive experiments on multiple benchmarks demonstrate the superiority of our framework. Without fine-tuning on novel classes, it brings significant performance gains to one-stage baselines, lifting state-of-the-art results to a higher level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
- When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object DetectionHongliang Zhou, Yongxiang Liu, Canyu Mo, Weijie Li et al.ICCV 2025 · 3 citations
- Exploring Base-Class Suppression with Prior Guidance for Bias-Free One-Shot Object DetectionWenwen Zhang, Yun Hu, Hangguan Shan, Eryun LiuAAAI 2024 · 3 citations
- PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationBing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang et al.ICCV 2025 · 1 citation
- BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram AlignmentRunmin Jiang, Jackson Daggett, Shriya Pingulkar, Yizhou Zhao et al.CVPR 2025
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
Related papers
- Adaptive Image Transformer for One-Shot Object DetectionDing-Jie Chen, He-Yen Hsieh, Tyng-Luh LiuCVPR 2021
- SpatialFormer: Semantic and Target Aware Attentions for Few-Shot LearningJinxiang Lai, Siqian Yang, Wenlong Wu, Tao Wu et al.AAAI 2023 · 21 citations
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen et al.CVPR 2022 · 183 citations
- CSDA: Learning Category-Scale Joint Feature for Domain Adaptive Object DetectionChanglong Gao, Chengxu Liu, Yujie Dun, Xueming QianICCV 2023 · 25 citations
- Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image ClassificationYike Wu, Bo Zhang, Gang Yu, Weixi Zhang et al.ACM MM 2021 · 40 citations
