LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei-Shi Zheng
Abstract
Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector cotraining with a large language model by generating imagelevel detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an imagelevel detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https: //github.com/iSEE-Laboratory/LLMDet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e905ce01-a5e1-4a90-bce9-800d0b426452Cited by top-tier papers8
- VL-SAM-V2: Open-World Object Detection with General and Specific Query FusionZhiwei Lin, Yongtao WangNeurIPS 2025 · 6 citations
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object DetectionSiheng Wang, Yanshu Li, Bohan Hu, Zhengdao Li et al.ICLR 2026 · 5 citations
- Prompt-Free Universal Region Proposal NetworkQihong Tang, Changhan Liu, Shaofeng Zhang, Wenbin Li et al.CVPR 2026 · 1 citation
- CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object DetectionZhichao Sun, Huazhang Hu, Yidong Ma, Gang Liu et al.NeurIPS 2025
- Fantastic Tractor-Dogs and How Not to Find Them With Open-Vocabulary DetectorsFrank Ruis, Gertjan J. Burghouts, Hugo KuijfICLR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object DetectionYasiru Ranasinghe, Elim Schenck, Florence Yellin, Shuowen Hu et al.CVPR 2026
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationCihang Peng, Qiming Hou, Zhong Ren, Kun ZhouICCV 2025
- DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object DetectionLewei Yao, Renjie Pi, Jianhua Han, Xiaodan Liang et al.CVPR 2024 · 22 citations
- CapDet: Unifying Dense Captioning and Open-World Detection PretrainingYanxin Long, Youpeng Wen, Jianhua Han, Hang Xu et al.CVPR 2023
- ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsHeng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding et al.CVPR 2025
