Object-level Semantic and Spatial Distillation for Open Vocabulary Detection
Zitong Li, Jinzhuo Wu, Fukang Zhao, Xinyue Wang, Jun Chen-CUG, Zhuo Cheng, Dapeng Luo
Abstract
Recent Open-vocabulary Object Detection (OVD) approaches adapt CLIP through region-level distillation to improve semantic alignment for novel categories. However, the distilled regional features are often used for both classification and localization, enhancing semantic consistency at the expense of spatial fidelity. To resolve this, we propose Object-level Semantic and Spatial Distillation (OSSD), a two-stage framework that explicitly decouples semantic and spatial feature learning. OSSD first distills object-level semantics from CLIP’s global [CLS] embeddings to enhance region discrimination, and then injects fine-grained spatial and structural priors via spatial distillation from a detector trained only on COCO base categories. Furthermore, we propose a Location Quality Estimation Head (LQEH) that predicts class-agnostic localization quality, complementing objectness confidence to improve the novel-object perception. Extensive experiments show that our method achieves 49.2 AP50 on the OV-COCO benchmark. exceeding the best previous result by 3.6%, On the OV-LVIS benchmark, our method reaches 40.5 mAP on novel categories, outperforming previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3319dca5-9a81-4bbf-a183-1c96089a8453Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan et al.CVPR 2022 · 209 citations
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan et al.NeurIPS 2023 · 138 citations
Related papers
- CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingXiaoshi Wu, Feng Zhu, Rui Zhao, Hongsheng LiCVPR 2023
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan et al.ACM MM 2024 · 5 citations
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan et al.NeurIPS 2022 · 215 citations
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan et al.ICCV 2023 · 35 citations
