Few-Shot Object Detection with Foundation Models
Guangxing Han, Ser-Nam Lim
Abstract
Few-shot object detection (FSOD) aims to detect objects with only a few training examples. Visual feature extraction and query-support similarity learning are the two critical components. Existing works are usually developed based on ImageNet pre-trained vision backbones and design sophisticated metric-learning networks for few-shot learning, but still have inferior accuracy. In this work, we study few-shot object detection using modern foundation models. First, vision-only contrastive pre-trained DINOv2 model is used for the vision backbone, which shows strong transferable performance without tuning the parameters. Second, Large Language Model (LLM) is employed for contextualized fewshot learning with the input of all classes and query image proposals. Language instructions are carefully designed to prompt the LLM to classify each proposal in context. The contextual information include proposal-proposal relations, proposal-class relations, and class-class relations, which can largely promote few-shot learning. We comprehensively evaluate the proposed model (FM-FSOD) in multiple FSOD benchmarks, achieving state-of-the-arts performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 851d307e-0757-4e13-a72d-ee60bd45fd7fCited by top-tier papers11
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningJianyang Gu, Sam Stevens, Elizabeth G. Campolongo, Matthew J. Thompson et al.NeurIPS 2025 · 60 citations
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li et al.NeurIPS 2025 · 8 citations
- FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph DiffusionChen-Bin Feng, Youyang Sha, Longfei Liu, Yongjun Yu et al.ICLR 2026 · 3 citations
- Visual Textualization for Image Prompted Object DetectionYongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei et al.ICCV 2025 · 1 citation
- FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy OptimizationMohammed Asad Karim, Vinay Kumar VermaICML 2026 · 1 citation
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Multi-modal Prototype Guided Few-shot Object DetectionChenbo Zhang, Bing Huangfu, Hongxu Ma, Jihong Guan et al.ACM MM 2025 · 3 citations
- Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature AlignmentGuangxing Han, Shiyuan Huang, Jiawei Ma, Yicheng He et al.AAAI 2022 · 227 citations
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen et al.CVPR 2022 · 183 citations
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- Accurate Few-Shot Object Detection With Support-Query Mutual Guidance and Hybrid LossLu Zhang, Shuigeng Zhou, Jihong Guan, Ji ZhangCVPR 2021
