LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
Sheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu, Shijian Lu
Abstract
Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge into detector training. However, most existing open-vocabulary detectors learn by aligning region embeddings with categorical labels (e.g., bicycle) only, disregarding the capability of VLMs on aligning visual embeddings with fine-grained text description of object parts (e.g., pedals and bells). This paper presents DVDet, a Descriptor-Enhanced Open Vocabulary Detector that introduces conditional context prompts and hierarchical textual descriptors that enable precise region-text alignment as well as open-vocabulary detection training in general. Specifically, the conditional context prompt transforms regional embeddings into image-like representations that can be directly integrated into general open vocabulary detection training. In addition, we introduce large language models as an interactive and implicit knowledge repository which enables iterative mining and refining visually oriented textual descriptors for precise region-text alignment. Extensive experiments over multiple large-scale benchmarks show that DVDet outperforms the state-of-the-art consistently by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6299f887-e907-4bcf-897a-76cf3b0a808eCited by top-tier papers19
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 42 citations
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao et al.NeurIPS 2024 · 31 citations
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng et al.ICCV 2025 · 6 citations
- Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation MatchingZhaoyang Li, Yuan Wang, Guoxin Xiong, Wangkai Li et al.ICCV 2025 · 5 citations
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object DetectionSiheng Wang, Yanshu Li, Bohan Hu, Zhengdao Li et al.ICLR 2026 · 5 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context ReasoningYanqi Li, Jianwei Niu, Ningbo Gu, Tao RenAAAI 2026
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao et al.CVPR 2024 · 8 citations
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 69 citations
- Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object DetectionYoujun Zhao, Jiaying Lin, Rynson W. H. LauAAAI 2025 · 2 citations
