Lune

NeurIPS2025顶会

DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection

Chiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello, Simone Calderara, Rita Cucchiara

2025年份
4被引次数
1顶会引用

摘要

Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing their abilities on multiple specialized domains remains essential. While recent methods rely on monolithic adaptation strategies with a single set of weights, we embrace modular deep learning. We introduce DitHub, a framework designed to build and maintain a library of efficient adaptation modules. Inspired by Version Control Systems, DitHub manages expert modules as branches that can be fetched and merged as needed. This modular approach allows us to conduct an in-depth exploration of the compositional properties of adaptation modules, marking the first such study in Object Detection. Our method achieves state-of-theart performance on the ODinW-13 benchmark and ODinW-O, a newly introduced benchmark designed to assess class reappearance.

Recent approaches, such as those detailed in [4], have extended Vision-Language detectors to incrementally accommodate new categories while preserving robust capabilities [40]. However, these methods employ monolithic adaptation, where all newly acquired knowledge is condensed into a single set of weights. This approach presents challenges in real-world scenarios that demand updates to specific concepts which may reappear under varied input modalities. For instance, in safety and security applications, the class "person" may need to be detected not only in standard RGB imagery but also in thermal imagery, necessitating continuous model adaptation to ensure consistent performance across different data types. Additionally, for rare or complex categories, incremental adaptation is crucial to refine the base model for fine-grained concepts. In such scenarios, a monolithic deep architecture faces challenges similar to managing a complex program written on 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

• We achieve state-of-the-art performance on ODinW-13 and the newly introduced ODinW-O.

• We conduct an in-depth analysis of the compositional and specialization capabilities of class-specific modules, a first in the context of Object Detection.

2 Related Works Incremental Vision-Language Object Detection. Vision-Language Object Detection (VLOD), also referred to as Open-Vocabulary detection [4], is the task of detecting and recognizing objects beyond a fixed set of predefined categories. Early approaches leverage powerful visual and textual encoders (e.g., CLIP [40]) to enable object detectors to recognize virtually any object specified by a textual query [10, 55]. A fundamental trend in this field has been the reliance on large-scale pre-training on extensive datasets [43], which has become standard practice for tasks such as classification. In the context of Vision-Language Object Detection, state-of-the-art methods have followed this paradigm, with GLIP [20] being one of the first notable examples. GLIP reframes Object Detection as a phrase-grounding problem, leading to a pre-trained model capable of generalizing to unseen objects by integrating textual and visual semantics from large-scale datasets. Building on GLIP's problem formulation, Grounding DINO [25] extends this paradigm by incorporating merged visual and textual semantics at multiple network stages. Recently, Incremental Vision-Language Object Detection (IVLOD) [4] has emerged as a natural extension of the VLOD paradigm. By incorporating Incremental Learning [47, 17, 42, 51], IVLOD addresses the challenge of fine-tuning a pre-trained Vision-Language model on specific object categories while preserving its zero-shot [35] capabilities. Following [4], we adopt Grounding DINO as the backbone and use IVLOD to explore the merging of class-specific modules.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper31

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖