Multi-modal Prototype Guided Few-shot Object Detection
Chenbo Zhang, Bing Huangfu, Hongxu Ma, Jihong Guan, Shuigeng Zhou
Abstract
Few-shot object detection (FSOD) is an important problem in computer vision, aiming to accurately detect objects with only a few annotated examples. Prototype learning has been widely explored in this field. Some methods extract visual prototypes from support images, but the limited sample size often leads to unrepresentative features. Others use textual prototypes generated by pre-trained vision-language models such as CLIP, which lack visual detail and may suffer from language ambiguity. As visual and textual prototypes offer complementary strengths --- detail and generalization respectively, single-modal prototypes struggle to balance both. To address this issue, we propose MP-DETR, the first multi-modal prototype guided method for FSOD. We design an adaptive multi-modal prototype fusion module to combine visual and textual prototypes from foundation models using a gating mechanism, producing multi-modal class prototypes that retain both general semantics and visual specificity. These prototypes are then deeply integrated into the DETR detection pipeline for guiding potential region selection, enhancing corresponding object queries, and constructing a prototype similarity-based classifier enhanced by contrastive learning to improve discrimination among similar classes. By incorporating multi-modal guidance into the detection process, MP-DETR achieves better performance than existing single-modal methods. Extensive experiments on MS-COCO and Pascal-VOC show that MP-DETR achieves SOTA results in various few-shot settings, confirming its effectiveness and superiority.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fc7511cc-0d2c-43ea-9214-bebac78fdfc8Cited by top-tier papers4
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou et al.KDD 2026 · 1 citation
- FlowTime: Towards Continuous Generative Watch Time Prediction via Flow-based Personalized PriorsHongxu Ma, Han Zhou, Chenghou Jin, Jie Zhang et al.KDD 2026 · 1 citation
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang et al.ICLR 2026
- PRISM: Progressive Robust Learning for Open-World Continual Category DiscoveryWei Feng, Sijin Zhou, Yiwen Jiang, Zongyuan GeICLR 2026
Related papers
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen et al.CVPR 2024 · 27 citations
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
- Few-Shot Object Detection with Foundation ModelsGuangxing Han, Ser-Nam LimCVPR 2024
