LMSeg: Language-guided Multi-dataset Segmentation
Qiang Zhou, Yuang Liu, Chaohui Yu, Jingliang Li, Zhibin Wang, Fan Wang
Abstract
It's a meaningful and attractive topic to build a general and inclusive segmentation model that can recognize more categories in various scenarios. A straightforward way is to combine the existing fragmented segmentation datasets and train a multi-dataset network. However, there are two major issues with multi-dataset segmentation: (1) the inconsistent taxonomy demands manual reconciliation to construct a unified taxonomy; (2) the inflexible one-hot common taxonomy causes time-consuming model retraining and defective supervision of unlabeled categories. In this paper, we investigate the multi-dataset segmentation and propose a scalable Language-guided Multi-dataset Segmentation framework, dubbed LMSeg, which supports both semantic and panoptic segmentation. Specifically, we introduce a pre-trained text encoder to map the category names to a text embedding space as a unified taxonomy, instead of using inflexible one-hot label. The model dynamically aligns the segment queries with the category embeddings. Instead of relabeling each dataset with the unified taxonomy, a category-guided decoding module is designed to dynamically guide predictions to each datasets taxonomy. Furthermore, we adopt a dataset-aware augmentation strategy that assigns each dataset a specific image augmentation pipeline, which can suit the properties of images from different datasets. Extensive experiments demonstrate that our method achieves significant improvements on four semantic and three panoptic segmentation datasets, and the ablation study evaluates the effectiveness of each component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ddf04b4-d2db-4fc7-8619-79fe54ea0aaaCited by top-tier papers7
- DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelXiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan et al.NeurIPS 2023 · 40 citations
- AIMS: All-Inclusive Multi-Level Segmentation for AnythingLu Qi, Jason Kuen, Weidong Guo, Jiuxiang Gu et al.NeurIPS 2023 · 9 citations
- Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNsRong Ma, Jie Chen, Xiangyang Xue, Jian PuNeurIPS 2024 · 3 citations
- Scaling up Image Segmentation across Data and TasksPei Wang, Zhaowei Cai, Hao Yang, Ashwin Swaminathan et al.CVPR 2025
- Mind the Gap: Transferring Labels to Align Object Detection DatasetsMikhail Kennerley, Angelica I Aviles-Rivero, Carola-Bibiane Schönlieb, Robby T. TanCVPR 2026
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
- Universal Segmentation at Arbitrary Granularity with Language InstructionYong Liu, Cairong Zhang, Yitong Wang, Jiahao Wang et al.CVPR 2024 · 15 citations
- MSeg: A Composite Dataset for Multi-Domain Semantic SegmentationJohn Lambert, Zhuang Liu, Ozan Sener, James Hays et al.CVPR 2020
- Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language EmbeddingLingchen Meng, Xiyang Dai, Yinpeng Chen, Pengchuan Zhang et al.CVPR 2023
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu et al.ICCV 2023 · 241 citations
