GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence Distillation
Xu Wang, Zilei Wang, Zihan Lin
Abstract
Incremental object detection (IOD) is a challenging task that requires detection models to continuously learn from newly arriving data. This work focuses on incremental learning for vision-language detectors (VLDs), an under explored domain. Existing research typically adopts a local alignment paradigm to avoid label conflicts, where different tasks are learned separately without interaction. However, we reveal that this practice fails to effectively preserve the semantic structure. Specifically, aligned relationships between objects and texts would collapse when handling novel categories, ultimately leading to catastrophic forgetting. Though knowledge distillation (KD) is a common approach for tackling this, traditional KD performs poorly when directly applied to VLDs, as for different phases, a natural knowledge gap exists in both encoding and decoding processes. To address above issues, we propose a novel method called Global alignment and Correspondence Distillation (GCD). Differently, we first integrate knowledge across phases within the same embedding space to construct global semantic structure. We then enable effective knowledge distillation in VLDs through a semantic correspondence mechanism, ensuring consistent proposal generation and decoding. On the top of that, we distill teacher model’s informative predictions and topological relationships to maintain stable local semantic structure. Extensive experiments on COCO 2017 demonstrate that our method significantly outperforms existing approaches, achieving new state-of-the-art in various IOD scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Parameterized Prompt for Incremental Object DetectionZijia An, Boyu Diao, Ruiqi Liu, Libo Huang et al.CVPR 2026 · 1 citation
- Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object DetectionAoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong et al.ICML 2026 · 1 citation
- YOLO-IOD: Towards Real Time Incremental Object DetectionShizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu et al.AAAI 2026 · 1 citation
- Boosting Vision-Language Models Towards Cross-Domain Incremental Object DetectionXu Wang, Zihan Lin, Yixin Zhang, Zilei WangCVPR 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan et al.CVPR 2022 · 209 citations
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin et al.ICCV 2023 · 133 citations
Related papers
- Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionHongquan Zhang, Bin-Bin Gao, Yi Zeng, Xudong Tian et al.AAAI 2024 · 12 citations
- Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge DistillationMengxue Kang, Jinpeng Zhang, Jinming Zhang, Xiashuang Wang et al.ICCV 2023 · 23 citations
- Symbiosis-Inspired Knowledge Distillation for Incremental Object DetectionMingyue Zeng, De Cheng, Zhipeng Xu, Huaijie Wang et al.ICML 2026
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan et al.ICCV 2023 · 35 citations
- Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response DistillationTao Feng, Mang Wang, Hangjie YuanCVPR 2022 · 101 citations
