ScaleDet: A Scalable Multi-Dataset Object Detector
Yanbei Chen, Manchen Wang, Abhay Mittal, Zhenlin Xu, Paolo Favaro, Joseph Tighe, Davide Modolo
Abstract
Multi-dataset training provides a viable solution for exploiting heterogeneous large-scale datasets without extra annotation cost. In this work, we propose a scalable multidataset detector (ScaleDet) that can scale up its generalization across datasets when increasing the number of training datasets. Unlike existing multi-dataset learners that mostly rely on manual relabelling efforts or sophisticated optimizations to unify labels across datasets, we introduce a simple yet scalable formulation to derive a unified semantic label space for multi-dataset training. ScaleDet is trained by visual-textual alignment to learn the label assignment with label semantic similarities across datasets. Once trained, ScaleDet can generalize well on any given upstream and downstream datasets with seen and unseen classes. We conduct extensive experiments using LVIS, COCO, Objects365, OpenImages as upstream datasets, and 13 datasets from Object Detection in the Wild (ODinW) as downstream datasets. Our results show that ScaleDet achieves compelling strong model performance with an mAP of 50.7 on LVIS, 58.8 on COCO, 46.8 on Objects365, 76.2 on OpenImages, and 71.8 on ODinW, surpassing stateof-the-art detectors with the same backbone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt TrainingXiaoyang Wu, Zhuotao Tian, Xin Wen, Bohao Peng et al.CVPR 2024 · 39 citations
- Transferring Labels to Solve Annotation Mismatches Across Object Detection DatasetsYuan-Hong Liao, David Acuna, Rafid Mahmood, James Lucas et al.ICLR 2024 · 4 citations
- Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNsRong Ma, Jie Chen, Xiangyang Xue, Jian PuNeurIPS 2024 · 3 citations
- Hyperbolic Learning with Synthetic Captions for Open-World DetectionFanjie Kong, Yanbei Chen, Jiarui Cai, Davide ModoloCVPR 2024
- Open-Det: An Efficient Learning Framework for Open-Ended DetectionGuiping Cao, Tao Wang, Wenjian Huang, Xiangyuan Lan et al.ICML 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
- DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelXiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan et al.NeurIPS 2023 · 40 citations
- Detecting Everything in the Open World: Towards Universal Object DetectionZhenyu Wang, Yali Li, Xi Chen, Ser-Nam Lim et al.CVPR 2023
- Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object DetectionHeng Zhang, Qiuyu Zhao, Linyu Zheng, Hao Zeng et al.CVPR 2024 · 6 citations
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
