Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers
Dahun Kim, Anelia Angelova, Weicheng Kuo
Abstract
We present Region-aware Open-vocabulary Vision Transformers (RO-ViT) -a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretraining phase, we propose to randomly crop and resize regions of positional embeddings instead of using the whole image positional embeddings. This better matches the use of positional embeddings at region-level in the detection finetuning phase. In addition, we replace the common softmax cross entropy loss in contrastive learning with focal loss to better learn the informative yet difficult examples. Finally, we leverage recent advances in novel object proposals to improve open-vocabulary detection finetuning. We evaluate our full model on the LVIS and COCO open-vocabulary detection benchmarks and zero-shot transfer. RO-ViT achieves a state-of-the-art 34.1 AP r on LVIS, surpassing the best existing approach by +7.8 points in addition to competitive zero-shot transfer detection. Surprisingly, RO-ViT improves the image-level representation as well and achieves the state of the art on 9 out of 12 metrics on COCO and Flickr imagetext retrieval benchmarks, outperforming competitive approaches with larger models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 460990e5-cffe-41a3-abda-36db553f0502Cited by top-tier papers54
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICLR 2024 · 129 citations
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
- Open-Vocabulary Video Anomaly DetectionPeng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun et al.CVPR 2024 · 56 citations
- Exploring Unbiased Deepfake Detection via Token-Level Shuffling and MixingXinghe Fu, Zhiyuan Yan, Taiping Yao, Shen Chen et al.AAAI 2025 · 41 citations
- Contrastive Feature Masking Open-Vocabulary Vision TransformerDahun Kim, Anelia Angelova, Weicheng KuoICCV 2023 · 40 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- CLIM: Contrastive Language-Image Mosaic for Region RepresentationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.AAAI 2024 · 30 citations
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic et al.ICCV 2023 · 22 citations
- Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsWeicheng Kuo, Yin Cui, Xiuye Gu, A. J. Piergiovanni et al.ICLR 2023 · 37 citations
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan et al.NeurIPS 2022 · 215 citations
