Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
Jieren Deng, Haojian Zhang, Kun Ding, Jianhua Hu, Xingxuan Zhang, Yunkuan Wang
摘要
This paper presents Incremental Vision-Language Object Detection (IVLOD), a novel learning task designed to incrementally adapt pre-trained Vision-Language Object Detection Models (VLODMs) to various specialized domains, while simultaneously preserving their zero-shot generalization capabilities for the generalized domain. To address this new challenge, we present the Zero-interference Reparameterizable Adaptation (ZiRa), a novel method that introduces Zero-interference Loss and reparameterization techniques to tackle IVLOD without incurring additional inference costs or a significant increase in memory usage. Comprehensive experiments on COCO and ODinW-13 datasets demonstrate that ZiRa effectively safeguards the zero-shot generalization ability of VLODMs while continuously adapting to new tasks. Specifically, after training on ODinW-13 datasets, ZiRa exhibits superior performance compared to CL-DETR and iDETR, boosting zero-shot generalizability by substantial 13.91 and 8.74 AP, respectively.Our code is available at https://github.com/JarintotionDin/ZiRaGroundingDINO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMBowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang 等NeurIPS 2025 · 被引用 25 次
- DitHub: A Modular Framework for Incremental Open-Vocabulary Object DetectionChiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello 等NeurIPS 2025 · 被引用 4 次
- EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformerMunish Monga, Vishal Chudasama, Pankaj Wasnik, C.V. JawaharCVPR 2026 · 被引用 1 次
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object DetectionWeihao Cao, Runqi Wang, Xiaoyue Duan, Jinchao Zhang 等CVPR 2026
- Boosting Vision-Language Models Towards Cross-Domain Incremental Object DetectionXu Wang, Zihan Lin, Yixin Zhang, Zilei WangCVPR 2026
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 被引用 845 次
相关 Paper
- Visual Modality Prompt for Adapting Vision-Language Object DetectorsHeitor Rapela Medeiros, Atif Belal, Srikanth Muralidharan, Eric Granger 等ICCV 2025 · 被引用 3 次
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationQinqian Lei, Bo Wang, Robby T. TanICCV 2025 · 被引用 4 次
- GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence DistillationXu Wang, Zilei Wang, Zihan LinAAAI 2025 · 被引用 4 次
- Learning Task-Aware Language-Image Representation for Class-Incremental Object DetectionHongquan Zhang, Bin-Bin Gao, Yi Zeng, Xudong Tian 等AAAI 2024 · 被引用 12 次
- Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningKun Ding, Haojian Zhang, Qiang Yu, Ying Wang 等AAAI 2024 · 被引用 8 次
