Lune

ACM MM2025顶会

Multi-modal Prototype Guided Few-shot Object Detection

Chenbo Zhang, Bing Huangfu, Hongxu Ma, Jihong Guan, Shuigeng Zhou

2025年份
3被引次数
4顶会引用

摘要

Few-shot object detection (FSOD) is an important problem in computer vision, aiming to accurately detect objects with only a few annotated examples. Prototype learning has been widely explored in this field. Some methods extract visual prototypes from support images, but the limited sample size often leads to unrepresentative features. Others use textual prototypes generated by pre-trained vision-language models such as CLIP, which lack visual detail and may suffer from language ambiguity. As visual and textual prototypes offer complementary strengths --- detail and generalization respectively, single-modal prototypes struggle to balance both. To address this issue, we propose MP-DETR, the first multi-modal prototype guided method for FSOD. We design an adaptive multi-modal prototype fusion module to combine visual and textual prototypes from foundation models using a gating mechanism, producing multi-modal class prototypes that retain both general semantics and visual specificity. These prototypes are then deeply integrated into the DETR detection pipeline for guiding potential region selection, enhancing corresponding object queries, and constructing a prototype similarity-based classifier enhanced by contrastive learning to improve discrimination among similar classes. By incorporating multi-modal guidance into the detection process, MP-DETR achieves better performance than existing single-modal methods. Extensive experiments on MS-COCO and Pascal-VOC show that MP-DETR achieves SOTA results in various few-shot settings, confirming its effectiveness and superiority.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖