FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
Chen-Bin Feng, Youyang Sha, Longfei Liu, Yongjun Yu, Chi-Man Vong, Xuanlong Yu, Xi Shen
Abstract
In this paper, we present FSOD-VFM: Few-Shot Object Detectors with Vision Foundation Models, a framework that leverages vision foundation models to tackle the challenge of few-shot object detection. FSOD-VFM integrates three key components: a universal proposal network (UPN) for category-agnostic bounding box generation, SAM2 for accurate mask extraction, and DINOv2 features for efficient adaptation to new object categories. Despite the strong generalization capabilities of foundation models, the bounding boxes generated by UPN often suffer from overfragmentation, covering only partial object regions and leading to numerous small, false-positive proposals rather than accurate, complete object detections. To address this issue, we introduce a novel graph-based confidence reweighting method. In our approach, predicted bounding boxes are modeled as nodes in a directed graph, with graph diffusion operations applied to propagate confidence scores across the network. This reweighting process refines the scores of proposals, assigning higher confidence to whole objects and lower confidence to local, fragmented parts. This strategy improves detection granularity and effectively reduces the occurrence of false-positive bounding box proposals. Through extensive experiments on Pascal-5, COCO-20, and CD-FSOD datasets, we demonstrate that our method substantially outperforms existing approaches, achieving superior performance without requiring additional training. Notably, on the challenging CD-FSOD dataset, which spans multiple datasets and domains, our FSOD-VFM achieves 31.6 AP in the 10-shot setting, substantially outperforming previous training-free methods that reach only 21.4 AP. Code is available at: https://intellindust-ai-lab.github.io/projects/FSOD-VFM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c19da117-232b-48eb-a90b-8d2dd0307a5aCited by top-tier papers1
Ask how each one uses itBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li et al.NeurIPS 2025 · 8 citations
- UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object DetectionYunhang Shen, Rongrong Ji, Zhiwei Chen, Yongjian Wu et al.NeurIPS 2020 · 37 citations
- Few-Shot Object Detection with Foundation ModelsGuangxing Han, Ser-Nam LimCVPR 2024
- Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature AlignmentGuangxing Han, Shiyuan Huang, Jiawei Ma, Yicheng He et al.AAAI 2022 · 227 citations
- TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object DetectionZhijin He, Shuo Jin, Siyue Yu, Shuwei Wu et al.CVPR 2026 · 1 citation
