CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object Detection
Shiyuan Ma, Donglin Qian, Kai Ye, Shengchuan Zhang
Abstract
Open vocabulary object detection (OVOD) task aims to detect objects of novel categories beyond the base categories in the training set. To this end, the detector needs to access image-text pairs containing rich semantic information or the visual language pre-trained model (VLM) learned on them. Recent OVOD methods rely on knowledge distillation from VLMs. However, there are two main problems in current methods: (1) Current knowledge distillation frameworks fail to take advantage of the global category information of VLMs and thus fail to learn category-specific knowledge. (2) Due to the overfitting phenomenon of base categories during training, current OVOD networks generally have the problem of suppressing novel categories as background. To address these two problems, we propose a Category Aware Knowledge Extraction framework (CAKE), which consists of a Category-Specific Knowledge Distillation branch (CSKD) and a Category Generalization Region Proposal Network (CG-RPN). CSKD can more fully extract category-strong related information through category-specific distillation, and it is also conducive to filtering the exclusion problem between individuals of the same category; in this process, the model constructs a category-specific feature set to maintain high-quality category features. CG-RPN leverages the guidance of feature set to adjust the confidence scores of region proposals, thereby mining proposals that potentially contain novel categories of objects. Extensive experiments show that our method can plug and play well with many existing methods and significantly improve their detection performance. Moreover, our CAKE framework can reach the-state-of-the-art performance on OV-COCO and OV-LVIS datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext deb5558b-917e-4ba0-8b5b-7a00cf9cc0bfCited by top-tier papers3
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object DetectionSiheng Wang, Yanshu Li, Bohan Hu, Zhengdao Li et al.ICLR 2026 · 5 citations
- SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object TrackingYangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao et al.AAAI 2026
- From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context ReasoningYanqi Li, Jianwei Niu, Ningbo Gu, Tao RenAAAI 2026
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan et al.ICCV 2023 · 35 citations
- Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object DetectionJiaming Li, Jiacheng Zhang, Jichang Li, Ge Li et al.CVPR 2024
- Tensor Decomposition and Language Description for Open-Vocabulary Object DetectionQiuyu Liang, Yongqiang ZhangAAAI 2026
- NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object DetectionYupeng Zhang, Ruize Han, Zhiwei Chen, Wei Feng et al.CVPR 2026 · 2 citations
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao et al.CVPR 2024 · 8 citations
