GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence Distillation
Xu Wang, Zilei Wang, Zihan Lin
摘要
Incremental object detection (IOD) is a challenging task that requires detection models to continuously learn from newly arriving data. This work focuses on incremental learning for vision-language detectors (VLDs), an under explored domain. Existing research typically adopts a local alignment paradigm to avoid label conflicts, where different tasks are learned separately without interaction. However, we reveal that this practice fails to effectively preserve the semantic structure. Specifically, aligned relationships between objects and texts would collapse when handling novel categories, ultimately leading to catastrophic forgetting. Though knowledge distillation (KD) is a common approach for tackling this, traditional KD performs poorly when directly applied to VLDs, as for different phases, a natural knowledge gap exists in both encoding and decoding processes. To address above issues, we propose a novel method called Global alignment and Correspondence Distillation (GCD). Differently, we first integrate knowledge across phases within the same embedding space to construct global semantic structure. We then enable effective knowledge distillation in VLDs through a semantic correspondence mechanism, ensuring consistent proposal generation and decoding. On the top of that, we distill teacher model’s informative predictions and topological relationships to maintain stable local semantic structure. Extensive experiments on COCO 2017 demonstrate that our method significantly outperforms existing approaches, achieving new state-of-the-art in various IOD scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Parameterized Prompt for Incremental Object DetectionZijia An, Boyu Diao, Ruiqi Liu, Libo Huang 等CVPR 2026 · 被引用 1 次
- Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object DetectionAoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong 等ICML 2026 · 被引用 1 次
- YOLO-IOD: Towards Real Time Incremental Object DetectionShizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu 等AAAI 2026 · 被引用 1 次
- Boosting Vision-Language Models Towards Cross-Domain Incremental Object DetectionXu Wang, Zihan Lin, Yixin Zhang, Zilei WangCVPR 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen 等NeurIPS 2020 · 被引用 2,118 次
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan 等CVPR 2022 · 被引用 209 次
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
相关 Paper
- Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionHongquan Zhang, Bin-Bin Gao, Yi Zeng, Xudong Tian 等AAAI 2024 · 被引用 12 次
- Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge DistillationMengxue Kang, Jinpeng Zhang, Jinming Zhang, Xiashuang Wang 等ICCV 2023 · 被引用 23 次
- Symbiosis-Inspired Knowledge Distillation for Incremental Object DetectionMingyue Zeng, De Cheng, Zhipeng Xu, Huaijie Wang 等ICML 2026
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan 等ICCV 2023 · 被引用 35 次
- Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response DistillationTao Feng, Mang Wang, Hangjie YuanCVPR 2022 · 被引用 101 次
