LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei-Shi Zheng
摘要
Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector cotraining with a large language model by generating imagelevel detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an imagelevel detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https: //github.com/iSEE-Laboratory/LLMDet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VL-SAM-V2: Open-World Object Detection with General and Specific Query FusionZhiwei Lin, Yongtao WangNeurIPS 2025 · 被引用 6 次
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object DetectionSiheng Wang, Yanshu Li, Bohan Hu, Zhengdao Li 等ICLR 2026 · 被引用 5 次
- Prompt-Free Universal Region Proposal NetworkQihong Tang, Changhan Liu, Shaofeng Zhang, Wenbin Li 等CVPR 2026 · 被引用 1 次
- CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object DetectionZhichao Sun, Huazhang Hu, Yidong Ma, Gang Liu 等NeurIPS 2025
- Fantastic Tractor-Dogs and How Not to Find Them With Open-Vocabulary DetectorsFrank Ruis, Gertjan J. Burghouts, Hugo KuijfICLR 2026
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object DetectionYasiru Ranasinghe, Elim Schenck, Florence Yellin, Shuowen Hu 等CVPR 2026
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationCihang Peng, Qiming Hou, Zhong Ren, Kun ZhouICCV 2025
- DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object DetectionLewei Yao, Renjie Pi, Jianhua Han, Xiaodan Liang 等CVPR 2024 · 被引用 22 次
- CapDet: Unifying Dense Captioning and Open-World Detection PretrainingYanxin Long, Youpeng Wen, Jianhua Han, Hang Xu 等CVPR 2023
- ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsHeng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding 等CVPR 2025
