Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
Xu Wang, Zihan Lin, Yixin Zhang, Zilei Wang
Abstract
Incremental Object Detection (IOD) aims to equip detectors with the ability to handle dynamic environments and emerging object categories, and the rise of vision-language models has substantially advanced this goal. However, existing studies often oversimplify real-world scenarios by assuming the incremental tasks come from a single general domain. To better investigate vision-language models under IOD, it is necessary to explore more generalized scenarios that encompass both novel categories and domains. To this end, we propose Cross-Domain Incremental Object Detection (CDIOD), a new benchmark that assesses the ability to continuously adapt to diverse object detection tasks across domains. CDIOD reveals that existing methods struggle to balance between adaptivity and stability under substantial domain shifts. To tackle this challenge, we propose Dynamic Group Subspace (DGS), a novel framework that dynamically groups tasks by distribution to promote knowledge sharing and prevent task collisions; progressively consolidates adapters to build shared subspaces and control parameter growth; and implements a dynamic training pipeline to maintain a proper stability-adaptivity balance. DGS enables vision-language models to effectively handle task streams of various distribution shifts. Extensive experiments across three benchmarks demonstrate that DGS achieves SOTA performance, highlighting its robustness in diverse incremental learning scenarios. Code is available at https://github.com/Never-wx/dgs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d7206ac-1a6e-4f1f-bfdf-b7a0430c8cb7Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- Zero-shot Generalizable Incremental Learning for Vision-Language Object DetectionJieren Deng, Haojian Zhang, Kun Ding, Jianhua Hu et al.NeurIPS 2024 · 22 citations
- DitHub: A Modular Framework for Incremental Open-Vocabulary Object DetectionChiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello et al.NeurIPS 2025 · 4 citations
- GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence DistillationXu Wang, Zilei Wang, Zihan LinAAAI 2025 · 4 citations
- AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object DetectionHaolin Li, Yaohua Wang, Ze Yan, Lijie Wen et al.CVPR 2026
- Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionHongquan Zhang, Bin-Bin Gao, Yi Zeng, Xudong Tian et al.AAAI 2024 · 12 citations
