Multi-modal Prototype Guided Few-shot Object Detection
Chenbo Zhang, Bing Huangfu, Hongxu Ma, Jihong Guan, Shuigeng Zhou
摘要
Few-shot object detection (FSOD) is an important problem in computer vision, aiming to accurately detect objects with only a few annotated examples. Prototype learning has been widely explored in this field. Some methods extract visual prototypes from support images, but the limited sample size often leads to unrepresentative features. Others use textual prototypes generated by pre-trained vision-language models such as CLIP, which lack visual detail and may suffer from language ambiguity. As visual and textual prototypes offer complementary strengths --- detail and generalization respectively, single-modal prototypes struggle to balance both. To address this issue, we propose MP-DETR, the first multi-modal prototype guided method for FSOD. We design an adaptive multi-modal prototype fusion module to combine visual and textual prototypes from foundation models using a gating mechanism, producing multi-modal class prototypes that retain both general semantics and visual specificity. These prototypes are then deeply integrated into the DETR detection pipeline for guiding potential region selection, enhancing corresponding object queries, and constructing a prototype similarity-based classifier enhanced by contrastive learning to improve discrimination among similar classes. By incorporating multi-modal guidance into the detection process, MP-DETR achieves better performance than existing single-modal methods. Extensive experiments on MS-COCO and Pascal-VOC show that MP-DETR achieves SOTA results in various few-shot settings, confirming its effectiveness and superiority.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou 等KDD 2026 · 被引用 1 次
- FlowTime: Towards Continuous Generative Watch Time Prediction via Flow-based Personalized PriorsHongxu Ma, Han Zhou, Chenghou Jin, Jie Zhang 等KDD 2026 · 被引用 1 次
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang 等ICLR 2026
- PRISM: Progressive Robust Learning for Open-World Continual Category DiscoveryWei Feng, Sijin Zhou, Yiwen Jiang, Zongyuan GeICLR 2026
相关 Paper
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke 等ICCV 2025 · 被引用 4 次
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen 等CVPR 2024 · 被引用 27 次
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 等AAAI 2024 · 被引用 71 次
- Few-Shot Object Detection with Foundation ModelsGuangxing Han, Ser-Nam LimCVPR 2024
