DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model
Xiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan, Xuan Yang, Xingyi Zhou, Golnaz Ghiasi, Weicheng Kuo, Huizhong Chen, Liang-Chieh Chen, David A. Ross
Abstract
Observing the close relationship among panoptic, semantic and instance segmentation tasks, we propose to train a universal multi-dataset multi-task segmentation model: DaTaSeg. We use a shared representation (mask proposals with class predictions) for all tasks. To tackle task discrepancy, we adopt different merge operations and post-processing for different tasks. We also leverage weak-supervision, allowing our segmentation model to benefit from cheaper bounding box annotations. To share knowledge across datasets, we use text embeddings from the same semantic embedding space as classifiers and share all network parameters among datasets. We train DaTaSeg on ADE semantic, COCO panoptic, and Objects365 detection datasets. DaTaSeg improves performance on all datasets, especially small-scale datasets, achieving 54.0 mIoU on ADE semantic and 53.5 PQ on COCO panoptic. DaTaSeg also enables weakly-supervised knowledge transfer on ADE panoptic and Objects365 instance segmentation. Experiments show DaTaSeg scales with the number of training datasets and enables open-vocabulary segmentation through direct transfer. In addition, we annotate an Objects365 instance segmentation set of 1,000 images and will release it as a public benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2023 · 285 citations
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- Describe Anything: Detailed Localized Image and Video CaptioningLong Lian, Yifan Ding, Yunhao Ge, Sifei Liu et al.ICCV 2025 · 14 citations
- One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object DetectionZhenyu Wang, Yali Li, Hengshuang Zhao, Shengjin WangNeurIPS 2024 · 13 citations
- Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language ModelsCong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng et al.ICCV 2025 · 8 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 500 citations
Related papers
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu et al.ICCV 2023 · 241 citations
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- OMG-Seg: Is One Model Good Enough for all Segmentation?Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding et al.CVPR 2024
- OneFormer: One Transformer to Rule Universal Image SegmentationJitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani et al.CVPR 2023
- Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2023
