LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
Sheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu, Shijian Lu
摘要
Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge into detector training. However, most existing open-vocabulary detectors learn by aligning region embeddings with categorical labels (e.g., bicycle) only, disregarding the capability of VLMs on aligning visual embeddings with fine-grained text description of object parts (e.g., pedals and bells). This paper presents DVDet, a Descriptor-Enhanced Open Vocabulary Detector that introduces conditional context prompts and hierarchical textual descriptors that enable precise region-text alignment as well as open-vocabulary detection training in general. Specifically, the conditional context prompt transforms regional embeddings into image-like representations that can be directly integrated into general open vocabulary detection training. In addition, we introduce large language models as an interactive and implicit knowledge repository which enables iterative mining and refining visually oriented textual descriptors for precise region-text alignment. Extensive experiments over multiple large-scale benchmarks show that DVDet outperforms the state-of-the-art consistently by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 被引用 42 次
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 31 次
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng 等ICCV 2025 · 被引用 6 次
- Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation MatchingZhaoyang Li, Yuan Wang, Guoxin Xiong, Wangkai Li 等ICCV 2025 · 被引用 5 次
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object DetectionSiheng Wang, Yanshu Li, Bohan Hu, Zhengdao Li 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
相关 Paper
- From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context ReasoningYanqi Li, Jianwei Niu, Ningbo Gu, Tao RenAAAI 2026
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao 等CVPR 2024 · 被引用 8 次
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi 等CVPR 2022 · 被引用 311 次
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 被引用 69 次
- Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object DetectionYoujun Zhao, Jiaying Lin, Rynson W. H. LauAAAI 2025 · 被引用 2 次
