Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding
Lingchen Meng, Xiyang Dai, Yinpeng Chen, Pengchuan Zhang, Dongdong Chen, Mengchen Liu, Jianfeng Wang, Zuxuan Wu, Lu Yuan, Yu-Gang Jiang
Abstract
Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy difference and domain gap. In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned. It not only mitigates the dataset inconsistency but also provides coherent guidance for the detector to learn across multiple datasets. In particular, the dataset-aware design is achieved by learning a dataset embedding that is used to adapt object queries as well as convolutional kernels in detection heads. The categories across datasets are semantically aligned into a unified space by replacing one-hot category representations with word embedding and leveraging the semantic coherence of language embedding. Detection Hub fulfills the benefits of large data on object detection. Experiments demonstrate that joint training on multiple datasets achieves significant performance gains over training on each dataset alone. Detection Hub further achieves SoTA performance on UODB benchmark with wide variety of datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionHanqing Zhao, Dianmo Sheng, Jianmin Bao, Dongdong Chen et al.ICML 2023 · 67 citations
- Multi-Prompt Alignment for Multi-Source Unsupervised Domain AdaptationHaoran Chen, Xintong Han, Zuxuan Wu, Yu-Gang JiangNeurIPS 2023 · 55 citations
- DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelXiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan et al.NeurIPS 2023 · 40 citations
- Learning from Rich Semantics and Coarse Locations for Long-tailed Object DetectionLingchen Meng, Xiyang Dai, Jianwei Yang, Dongdong Chen et al.NeurIPS 2023 · 23 citations
- One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object DetectionZhenyu Wang, Yali Li, Hengshuang Zhao, Shengjin WangNeurIPS 2024 · 13 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
Related papers
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
- Uni2Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D DetectionYubin Wang, Zhikang Zou, Xiaoqing Ye, Xiao Tan et al.ICLR 2025
- LMSeg: Language-guided Multi-dataset SegmentationQiang Zhou, Yuang Liu, Chaohui Yu, Jingliang Li et al.ICLR 2023 · 2 citations
- CSDA: Learning Category-Scale Joint Feature for Domain Adaptive Object DetectionChanglong Gao, Chengxu Liu, Yujie Dun, Xueming QianICCV 2023 · 25 citations
- Mind the Gap: Transferring Labels to Align Object Detection DatasetsMikhail Kennerley, Angelica I Aviles-Rivero, Carola-Bibiane Schönlieb, Robby T. TanCVPR 2026
