FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-training
Adrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios Tzimiropoulos
摘要
This paper is on Few-Shot Object Detection (FSOD), where given a few templates (examples) depicting a novel class (not seen during training), the goal is to detect all of its occurrences within a set of images. From a practical perspective, an FSOD system must fulfil the following desiderata: (a) it must be used as is, without requiring any fine-tuning at test time, (b) it must be able to process an arbitrary number of novel objects concurrently while supporting an arbitrary number of examples from each class and (c) it must achieve accuracy comparable to a closed system. Towards satisfying (a)-(c), in this work, we make the following contributions: We introduce, for the first time, a simple, yet powerful, few-shot detection transformer (FS-DETR) based on visual prompting that can address both desiderata (a) and (b). Our system builds upon the DETR framework, extending it based on two key ideas: (1) feed the provided visual templates of the novel classes as visual prompts during test time, and (2) "stamp" these prompts with pseudo-class embeddings (akin to soft prompting), which are then predicted at the output of the decoder. Importantly, we show that our system is not only more flexible than existing methods, but also, it makes a step towards satisfying desideratum (c). Specifically, it is significantly more accurate than all methods that do not require fine-tuning and even matches and outperforms the current state-of-the-art fine-tuning based methods on the most well-established benchmarks (PASCAL VOC & MSCOCO).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling AugmentationYanjie Wang, Xu Zou, Luxin Yan, Sheng Zhong 等CVPR 2024 · 被引用 22 次
- Towards Single-Source Domain Generalized Object Detection via Causal Visual PromptsChen Li, Huiying Xu, Changxin Gao, Zeyu Wang 等NeurIPS 2025 · 被引用 3 次
- Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View ImagesHongyu Yan, Yadong MuAAAI 2025 · 被引用 3 次
- Visual Textualization for Image Prompted Object DetectionYongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei 等ICCV 2025 · 被引用 1 次
- Few-Shot Object Detection with Foundation ModelsGuangxing Han, Ser-Nam LimCVPR 2024
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Weakly Supervised Few-Shot Object Detection with DETRChenbo Zhang, Yinglu Zhang, Lu Zhang, Jiajia Zhao 等AAAI 2024 · 被引用 8 次
- Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised LearningNa Dong, Yongqiang Zhang, Mingli Ding, Gim Hee LeeAAAI 2023 · 被引用 54 次
- PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object DetectionYingjie Gao, Yanan Zhang, Ziyue Huang, Nanqing Liu 等ACM MM 2024 · 被引用 13 次
- Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale ApproachMir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. LittleCVPR 2024
- Label, Verify, Correct: A Simple Few Shot Object Detection MethodPrannay Kaul, Weidi Xie, Andrew ZissermanCVPR 2022 · 被引用 123 次
