DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection
Chiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello, Simone Calderara, Rita Cucchiara
Abstract
Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing their abilities on multiple specialized domains remains essential. While recent methods rely on monolithic adaptation strategies with a single set of weights, we embrace modular deep learning. We introduce DitHub, a framework designed to build and maintain a library of efficient adaptation modules. Inspired by Version Control Systems, DitHub manages expert modules as branches that can be fetched and merged as needed. This modular approach allows us to conduct an in-depth exploration of the compositional properties of adaptation modules, marking the first such study in Object Detection. Our method achieves state-of-theart performance on the ODinW-13 benchmark and ODinW-O, a newly introduced benchmark designed to assess class reappearance.
Recent approaches, such as those detailed in [4], have extended Vision-Language detectors to incrementally accommodate new categories while preserving robust capabilities [40]. However, these methods employ monolithic adaptation, where all newly acquired knowledge is condensed into a single set of weights. This approach presents challenges in real-world scenarios that demand updates to specific concepts which may reappear under varied input modalities. For instance, in safety and security applications, the class "person" may need to be detected not only in standard RGB imagery but also in thermal imagery, necessitating continuous model adaptation to ensure consistent performance across different data types. Additionally, for rare or complex categories, incremental adaptation is crucial to refine the base model for fine-grained concepts. In such scenarios, a monolithic deep architecture faces challenges similar to managing a complex program written on 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
• We achieve state-of-the-art performance on ODinW-13 and the newly introduced ODinW-O.
• We conduct an in-depth analysis of the compositional and specialization capabilities of class-specific modules, a first in the context of Object Detection.
2 Related Works Incremental Vision-Language Object Detection. Vision-Language Object Detection (VLOD), also referred to as Open-Vocabulary detection [4], is the task of detecting and recognizing objects beyond a fixed set of predefined categories. Early approaches leverage powerful visual and textual encoders (e.g., CLIP [40]) to enable object detectors to recognize virtually any object specified by a textual query [10, 55]. A fundamental trend in this field has been the reliance on large-scale pre-training on extensive datasets [43], which has become standard practice for tasks such as classification. In the context of Vision-Language Object Detection, state-of-the-art methods have followed this paradigm, with GLIP [20] being one of the first notable examples. GLIP reframes Object Detection as a phrase-grounding problem, leading to a pre-trained model capable of generalizing to unseen objects by integrating textual and visual semantics from large-scale datasets. Building on GLIP's problem formulation, Grounding DINO [25] extends this paradigm by incorporating merged visual and textual semantics at multiple network stages. Recently, Incremental Vision-Language Object Detection (IVLOD) [4] has emerged as a natural extension of the VLOD paradigm. By incorporating Incremental Learning [47, 17, 42, 51], IVLOD addresses the challenge of fine-tuning a pre-trained Vision-Language model on specific object categories while preserving its zero-shot [35] capabilities. Following [4], we adopt Grounding DINO as the backbone and use IVLOD to explore the merging of class-specific modules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbd8b382-f906-4e4e-9da1-dea0395c56dfCited by top-tier papers1
Ask how each one uses itBuilds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- Zero-shot Generalizable Incremental Learning for Vision-Language Object DetectionJieren Deng, Haojian Zhang, Kun Ding, Jianhua Hu et al.NeurIPS 2024 · 22 citations
- Towards Universal Perception through Language-Guided Open-World Object DetectionZihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long et al.ACM MM 2025 · 1 citation
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 69 citations
- Streamlined Open-Vocabulary Human-Object Interaction DetectionChang Sun, Dongliang Liao, Changxing DingCVPR 2026 · 2 citations
- Learning to Detect and Segment for Open Vocabulary Object DetectionTao WangCVPR 2023
