Lune

NeurIPS2025Top-tier venue

DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection

Chiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello, Simone Calderara, Rita Cucchiara

2025Year
4Citations
1Top-tier citations

Abstract

Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing their abilities on multiple specialized domains remains essential. While recent methods rely on monolithic adaptation strategies with a single set of weights, we embrace modular deep learning. We introduce DitHub, a framework designed to build and maintain a library of efficient adaptation modules. Inspired by Version Control Systems, DitHub manages expert modules as branches that can be fetched and merged as needed. This modular approach allows us to conduct an in-depth exploration of the compositional properties of adaptation modules, marking the first such study in Object Detection. Our method achieves state-of-theart performance on the ODinW-13 benchmark and ODinW-O, a newly introduced benchmark designed to assess class reappearance.

Recent approaches, such as those detailed in [4], have extended Vision-Language detectors to incrementally accommodate new categories while preserving robust capabilities [40]. However, these methods employ monolithic adaptation, where all newly acquired knowledge is condensed into a single set of weights. This approach presents challenges in real-world scenarios that demand updates to specific concepts which may reappear under varied input modalities. For instance, in safety and security applications, the class "person" may need to be detected not only in standard RGB imagery but also in thermal imagery, necessitating continuous model adaptation to ensure consistent performance across different data types. Additionally, for rare or complex categories, incremental adaptation is crucial to refine the base model for fine-grained concepts. In such scenarios, a monolithic deep architecture faces challenges similar to managing a complex program written on 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

• We achieve state-of-the-art performance on ODinW-13 and the newly introduced ODinW-O.

• We conduct an in-depth analysis of the compositional and specialization capabilities of class-specific modules, a first in the context of Object Detection.

2 Related Works Incremental Vision-Language Object Detection. Vision-Language Object Detection (VLOD), also referred to as Open-Vocabulary detection [4], is the task of detecting and recognizing objects beyond a fixed set of predefined categories. Early approaches leverage powerful visual and textual encoders (e.g., CLIP [40]) to enable object detectors to recognize virtually any object specified by a textual query [10, 55]. A fundamental trend in this field has been the reliance on large-scale pre-training on extensive datasets [43], which has become standard practice for tasks such as classification. In the context of Vision-Language Object Detection, state-of-the-art methods have followed this paradigm, with GLIP [20] being one of the first notable examples. GLIP reframes Object Detection as a phrase-grounding problem, leading to a pre-trained model capable of generalizing to unseen objects by integrating textual and visual semantics from large-scale datasets. Building on GLIP's problem formulation, Grounding DINO [25] extends this paradigm by incorporating merged visual and textual semantics at multiple network stages. Recently, Incremental Vision-Language Object Detection (IVLOD) [4] has emerged as a natural extension of the VLOD paradigm. By incorporating Incremental Learning [47, 17, 42, 51], IVLOD addresses the challenge of fine-tuning a pre-trained Vision-Language model on specific object categories while preserving its zero-shot [35] capabilities. Following [4], we adopt Grounding DINO as the backbone and use IVLOD to explore the merging of class-specific modules.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dbd8b382-f906-4e4e-9da1-dea0395c56df

Cited by top-tier papers1

Ask how each one uses it

Builds on31

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines