DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment
Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Hang Xu
2023Year
48Top-tier citations
Abstract
Hand drawn black and white illustration, flying eagle with a snake in claws.
(c) Curator looks on as we consider paintings in a shared studio space (b) Fried egg on a frying pan with cherry tomatoes and parsley (e) Little girl witch with black cat, owl, the witch's cauldron, ghost spirits and text on violet background. The concept of Halloween.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers48
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
- CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object DetectionChuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan et al.NeurIPS 2023 · 88 citations
- Multi-modal Queried Object Detection in the WildYifan Xu, Mengdan Zhang, Chaoyou Fu, Peixian Chen et al.NeurIPS 2023 · 73 citations
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Exploring Sparse MoE in GANs for Text-conditioned Image SynthesisJiapeng Zhu, Ceyuan Yang, Kecheng Zheng, Yinghao Xu et al.CVPR 2025
- One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency ControlsMinghui Hu, Jianbin Zheng, Chuanxia Zheng, Chaoyue Wang et al.CVPR 2024
- Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion ModelsRonghuan Wu, Wanchao Su, Jing LiaoCVPR 2025
- HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingMude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi et al.ICLR 2025
- TransPixeler: Advancing Text-to-Video Generation with TransparencyLuozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang et al.CVPR 2025
