Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection
Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Danfeng Qin, Jesse Berent
Abstract
Learning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is naturally available in the form of unstructured textual descriptions that users may freely provide when uploading image content. However, straightforward approaches to using such data for WSOD wastefully discard captions that do not exactly match object names. Instead, we show how to squeeze the most information out of these captions by training a text-only classifier that generalizes beyond dataset boundaries. Our discovery provides an opportunity for learning detection models from noisy but more abundant and freely-available caption data. We also validate our model on three classic object detection benchmarks and achieve state-of-the-art WSOD performance. Our code is available at https://github. com/yekeren/Cap2Det .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a3520cf-9507-4eb9-8759-c8236bb1d903Cited by top-tier papers21
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan et al.NeurIPS 2022 · 215 citations
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig et al.NeurIPS 2023 · 93 citations
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal PretrainingXunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei et al.ICCV 2021 · 84 citations
Related papers
- Weakly Supervised Few-Shot Object Detection with DETRChenbo Zhang, Yinglu Zhang, Lu Zhang, Jiajia Zhao et al.AAAI 2024 · 8 citations
- UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object DetectionYunhang Shen, Rongrong Ji, Zhiwei Chen, Yongjian Wu et al.NeurIPS 2020 · 37 citations
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 29 citations
- Seeing the Whole Through the Parts: Discovering Objects through Semantic Part Mining in Weak SupervisionShucheng Li, Weixuan Xu, Le Jiang, Hao Wu et al.SIGIR 2026
- Open-Vocabulary Object Detection Using CaptionsAlireza Zareian, Kevin Dela Rosa, Derek Hao Hu, Shih-Fu ChangCVPR 2021
