Transformer-based Open-world Instance Segmentation with Cross-task Consistency Regularization
Xizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu, Satoshi Tsutsui, Ying Li, Zehuan Yuan, Ping Song, Mike Zheng Shou
Abstract
Open-World Instance Segmentation (OWIS) is an emerging research topic that aims to segment class-agnostic object instances from images. The mainstream approaches use a two-stage segmentation framework, which first locates the candidate object bounding boxes and then performs instance segmentation. In this work, we instead promote a single-stage transformer-based framework for OWIS. We argue that the end-to-end training process in the single-stage framework can be more convenient for directly regularizing the localization of class-agnostic object pixels. Based on the transformer-based instance segmentation framework, we propose a regularization model to predict foreground pixels and use its relation to instance segmentation to construct a cross-task consistency loss. We show that such a consistency loss could alleviate the problem of incomplete instance annotation - a common problem in the existing OWIS datasets. We also show that the proposed loss lends itself to an effective solution to semi-supervised OWIS that could be considered an extreme case that all object annotations are absent for some images. Our extensive experiments demonstrate that the proposed method achieves impressive results in both fully-supervised and semi-supervised settings. Compared to SOTA methods, the proposed method significantly improves the AP_100 score by 4.75% in UVO dataset →UVO dataset setting and 4.05% in COCO dataset →UVO dataset setting.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4f3ed156-bd8e-468a-9295-458ef4ebeed5Cited by top-tier papers1
Ask how each one uses itRelated papers
- Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2023
- FreeSOLO: Learning to Segment Objects without AnnotationsXinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz et al.CVPR 2022 · 100 citations
- UN-DETR: Promoting Objectness Learning via Joint Supervision for Unknown Object DetectionHaomiao Liu, Hao Xu, Chuhuai Yue, Bo MaAAAI 2025 · 1 citation
- SOIT: Segmenting Objects with Instance-Aware TransformersXiaodong Yu, Dahu Shi, Xing Wei, Ye Ren et al.AAAI 2022 · 32 citations
- Semi-DETR: Semi-Supervised Object Detection with Detection TransformersJiacheng Zhang, Xiangru Lin, Wei Zhang, Kuo Wang et al.CVPR 2023
