Lune

NeurIPS2023Top-tier venue

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan, Xiaojuan Qi

2023Year
88Citations
22Top-tier citations

Abstract

Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-language models for alignment, which are prone to limitations in localization accuracy or generalization capabilities. In this paper, we propose CoDet, a novel approach that overcomes the reliance on pre-aligned vision-language space by reformulating region-word alignment as a co-occurring object discovery problem. Intuitively, by grouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept. Extensive experiments demonstrate that CoDet has superior performances and compelling scalability in open-vocabulary detection, e.g., by scaling up the visual backbone, CoDet achieves 37.0 AP m novel and 44.7 AP m all on OV-LVIS, surpassing the previous SoTA by 4.2 AP m novel and 9.8 AP m all . Code is available at https://github.com/CVMI-Lab/CoDet . Recent studies typically rely on vision-language models (VLMs) to determine region-word alignments, for example, by estimating region-word similarity [59, 27, 15, 28] . Despite its simplicity, the quality of generated pseudo region-text pairs is subject to limitations of VLMs. As illustrated in Figure 1b , VLMs pre-trained with image-level supervision, such as CLIP [35] , are largely unaware of localization quality of pseudo labels [59] . 53] mitigate this issue to some extent, they are initially pre-trained with a limited number of detection or grounding concepts, † This work was performed when Chuofan Ma worked as an intern at ByteDance. * Equal contribution. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bf22e593-e6d9-478d-83e7-8844165c9200

Cited by top-tier papers22

Ask how each one uses it

Builds on31

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines