PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li, Yuhuan Lin, Hanqiu Deng, Wenbing Tao, Yong Liu, Chengjie Wang
摘要
Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text paired samples for rare categories. This results in suboptimal performance in specialized domains or with complex objects. Recent visual-prompted methods partially address these issues but often involve complex multi-modal designs and multi-stage optimizations, extending the development cycle. Additionally, effective training strategies for data-driven OSOD models remain largely unexplored. To address these challenges, we propose PET-DINO, a universal object detector supporting both text and visual prompts. Our visual prompt generation scheme builds on an advanced text-prompted detector, addressing the limitations of text representation guidance and reducing the development cycle. We introduce two prompt-enriched training strategies: Intra-Batch Parallel Prompting (IBP) at the iteration level and Dynamic Memory-Driven Prompting (DMD) at the overall training level. These strategies enable simultaneous modeling of multiple prompt routes, parallel alignment with diverse real-world usage scenarios, and improved classification. Extensive experiments demonstrate that our visual prompt generation scheme, based on text-prompt-based detection pretraining, achieves a higher performance ceiling compared to using visual prompts alone.Our method achieves significant zero-shot detection performance on COCO, LVIS, and ODinW, and excels across various prompt-based detection protocols. In-domain evaluations also demonstrate robust localization performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen 等NeurIPS 2020 · 被引用 2,118 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
相关 Paper
- Visual in-Context PromptingFeng Li, Qing Jiang, Hao Zhang, Tianhe Ren 等CVPR 2024
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi 等CVPR 2022 · 被引用 311 次
- Text-Guided Visual Prompt DINO for Generic SegmentationYuchen Guan, Chong Sun, Canmiao Fu, Zhipeng Huang 等ICCV 2025 · 被引用 3 次
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao 等CVPR 2024 · 被引用 8 次
- CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object DetectionQibo Chen, Weizhong Jin, Jianyue Ge, Mengdi Liu 等AAAI 2025 · 被引用 3 次
