Video Instance Segmentation
Linjie Yang, Yuchen Fan, Ning Xu
Abstract
In this paper we present a new computer vision task, named video instance segmentation. The goal of this new task is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain. To facilitate research on this new task, we propose a large-scale benchmark called YouTube-VIS, which consists of 2,883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks. In addition, we propose a novel algorithm called Mask-Track R-CNN for this task. Our new method introduces a new tracking branch to Mask R-CNN to jointly perform the detection, segmentation and tracking tasks simultaneously. Finally, we evaluate the proposed method and several strong baselines on our new dataset. Experimental results clearly demonstrate the advantages of the proposed algorithm and reveal insight for future improvement. We believe the video instance segmentation task will motivate the community along the line of research for video understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers181
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention NetworksVivien Sainte Fare Garnot, Loïc LandrieuICCV 2021 · 245 citations
Related papers
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
- End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksTao Wang, Ning Xu, Kean Chen, Weiyao LinICCV 2021 · 30 citations
- SG-Net: Spatial Granularity Network for One-Stage Video Instance SegmentationDongfang Liu, Yiming Cui, Wenbo Tan, Yingjie Victor ChenCVPR 2021
- Video Instance Segmentation with a Propose-Reduce ParadigmHuaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu et al.ICCV 2021 · 110 citations
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu et al.ICCV 2023 · 56 citations
