CapDet: Unifying Dense Captioning and Open-World Detection Pretraining
Yanxin Long, Youpeng Wen, Jianhua Han, Hang Xu, Pengzhen Ren, Wei Zhang, Shen Zhao, Xiaodan Liang
Abstract
Benefiting from large-scale vision-language pre-training on image-text pairs, open-world detection methods have shown superior generalization ability under the zero-shot or few-shot detection settings. However, a pre-defined category space is still required during the inference stage of existing methods and only the objects belonging to that space will be predicted. To introduce a "real" open-world detector, in this paper, we propose a novel method named CapDet to either predict under a given category list or directly generate the category of predicted bounding boxes. Specifically, we unify the open-world detection and dense caption tasks into a single yet effective framework by introducing an additional dense captioning head to generate the region-grounded captions. Besides, adding the captioning task will in turn benefit the generalization of detection performance since the captioning dataset covers more concepts. Experiment results show that by unifying the dense caption task, our CapDet has obtained significant performance improvements (e.g., +2.1% mAP on LVIS rare classes) over the baseline method on LVIS (1203 classes). Besides, our CapDet also achieves state-of-the-art performance on dense captioning tasks, e.g., 15.44% mAP on VG V1.2 and 13.98% on the VG-COCO dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59274077-87d6-4f2a-80d2-3e05fe43fa03Cited by top-tier papers18
- Multi-modal Queried Object Detection in the WildYifan Xu, Mengdan Zhang, Chaoyou Fu, Peixian Chen et al.NeurIPS 2023 · 73 citations
- YOLOE: Real-Time Seeing AnythiAo Wang, Lihao Liu, Hui Chen, Zijia Lin et al.ICCV 2025 · 52 citations
- DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object DetectionLewei Yao, Renjie Pi, Jianhua Han, Xiaodan Liang et al.CVPR 2024 · 22 citations
- WeDetect: Fast Open-Vocabulary Object Detection as RetrievalShenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu et al.CVPR 2026 · 11 citations
- OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and CaptioningAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingNeurIPS 2024 · 7 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Hyperbolic Learning with Synthetic Captions for Open-World DetectionFanjie Kong, Yanbei Chen, Jiarui Cai, Davide ModoloCVPR 2024
- Learning Object-Language Alignments for Open-Vocabulary Object DetectionChuang Lin, Peize Sun, Yi Jiang, Ping Luo et al.ICLR 2023 · 36 citations
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object DetectionHeng Zhang, Qiuyu Zhao, Linyu Zheng, Hao Zeng et al.CVPR 2024 · 6 citations
- Generative Region-Language Pretraining for Open-Ended Object DetectionChuang Lin, Yi Jiang, Lizhen Qu, Zehuan Yuan et al.CVPR 2024
