Video Panoptic Segmentation
Dahun Kim, Sanghyun Woo, Joon-Young Lee, In So Kweon
Abstract
Panoptic segmentation has become a new standard of visual recognition task by unifying previous semantic segmentation and instance segmentation tasks in concert. In this paper, we propose and explore a new video extension of this task, called video panoptic segmentation. The task requires generating consistent panoptic segmentation as well as an association of instance ids across video frames. To invigorate research on this new task, we present two types of video panoptic datasets. The first is a re-organization of the synthetic VIPER dataset into the video panoptic format to exploit its large-scale pixel annotations. The second is a temporal extension on the Cityscapes val. set, by providing new video panoptic annotations (Cityscapes-VPS). Moreover, we propose a novel video panoptic segmentation network (VPSNet) which jointly predicts object classes, bounding boxes, masks, instance id tracking, and semantic segmentation in video frames. To provide appropriate metrics for this task, we propose a video panoptic quality (VPQ) metric and evaluate our method and several other baselines. Experimental results demonstrate the effectiveness of the presented two datasets. We achieve state-of-the-art results in image PQ on Cityscapes and also in VPQ on Cityscapes-VPS and VIPER datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bb81614-6837-4684-af4a-367ec970541bCited by top-tier papers7
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- Universal Adversarial Perturbations Through the Lens of Deep Steganography: Towards a Fourier PerspectiveChaoning Zhang, Philipp Benz, Adil Karjauv, In So KweonAAAI 2021 · 50 citations
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia et al.NeurIPS 2022 · 23 citations
- Robust and Consistent Online Video Instance Segmentation via Instance Mask PropagationMiran Heo, Seoung Wug Oh, Seon Joo Kim, Joon-Young LeeAAAI 2025 · 2 citations
- Multi-Granularity Video Object SegmentationSangbeom Lim, Seongchan Kim, Seungjun An, Seokju Cho et al.AAAI 2025
Builds on5
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- AdaptIS: Adaptive Instance Selection NetworkKonstantin Sofiiuk, Olga Barinova, Anton KonushinICCV 2019 · 179 citations
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of ThingsCheng-Yang Fu, Tamara L. Berg, Alexander C. BergICCV 2019 · 18 citations
- Learning Instance Occlusion for Panoptic SegmentationJustin Lazarow, Kwonjoon Lee, Kunyu Shi, Zhuowen TuCVPR 2020
Related papers
- Slot-VPS: Object-centric Representation Learning for Video Panoptic SegmentationYi Zhou, Hui Zhang, Hana Lee, Shuyang Sun et al.CVPR 2022 · 20 citations
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li et al.CVPR 2022 · 58 citations
- Learning To Associate Every Segment for Video Panoptic SegmentationSanghyun Woo, Dahun Kim, Joon-Young Lee, In So KweonCVPR 2021
- TarViS: A Unified Approach for Target-Based Video SegmentationAli Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan et al.CVPR 2023
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
