Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
Xu Wang, Zihan Lin, Yixin Zhang, Zilei Wang
摘要
Incremental Object Detection (IOD) aims to equip detectors with the ability to handle dynamic environments and emerging object categories, and the rise of vision-language models has substantially advanced this goal. However, existing studies often oversimplify real-world scenarios by assuming the incremental tasks come from a single general domain. To better investigate vision-language models under IOD, it is necessary to explore more generalized scenarios that encompass both novel categories and domains. To this end, we propose Cross-Domain Incremental Object Detection (CDIOD), a new benchmark that assesses the ability to continuously adapt to diverse object detection tasks across domains. CDIOD reveals that existing methods struggle to balance between adaptivity and stability under substantial domain shifts. To tackle this challenge, we propose Dynamic Group Subspace (DGS), a novel framework that dynamically groups tasks by distribution to promote knowledge sharing and prevent task collisions; progressively consolidates adapters to build shared subspaces and control parameter growth; and implements a dynamic training pipeline to maintain a proper stability-adaptivity balance. DGS enables vision-language models to effectively handle task streams of various distribution shifts. Extensive experiments across three benchmarks demonstrate that DGS achieves SOTA performance, highlighting its robustness in diverse incremental learning scenarios. Code is available at https://github.com/Never-wx/dgs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
相关 Paper
- Zero-shot Generalizable Incremental Learning for Vision-Language Object DetectionJieren Deng, Haojian Zhang, Kun Ding, Jianhua Hu 等NeurIPS 2024 · 被引用 22 次
- DitHub: A Modular Framework for Incremental Open-Vocabulary Object DetectionChiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello 等NeurIPS 2025 · 被引用 4 次
- GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence DistillationXu Wang, Zilei Wang, Zihan LinAAAI 2025 · 被引用 4 次
- AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object DetectionHaolin Li, Yaohua Wang, Ze Yan, Lijie Wen 等CVPR 2026
- Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionHongquan Zhang, Bin-Bin Gao, Yi Zeng, Xudong Tian 等AAAI 2024 · 被引用 12 次
