Scene-Centric Unsupervised Video Panoptic Segmentation
Christoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe, Christian Rupprecht, Daniel Cremers, Stefan Roth
Abstract
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focused on image segmentation tasks; the video domain remains underexplored. We propose CUViPS, the first unsupervised VPS approach. CUViPS generates temporally consistent panoptic video pseudo-labels from monocular scene-centric videos by exploiting unsupervised depth, motion, and visual cues. Training on these pseudo-labels using a novel Video DropLoss yields an accurate and unsupervised VPS model. To benchmark progress, we introduce a comprehensive evaluation protocol and four competitive baselines, extending state-of-the-art unsupervised panoptic image and instance video segmentation models to VPS. CUViPS consistently outperforms all baselines and demonstrates strong label-efficient learning. With CUViPS, our evaluation protocol, and baselines, we provide a strong foundation for future research on unsupervised VPS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67005168-5592-4582-93ab-4b4ac2db5963Builds on51
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
Related papers
- Scene-Centric Unsupervised Panoptic SegmentationOliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers et al.CVPR 2025
- Slot-VPS: Object-centric Representation Learning for Video Panoptic SegmentationYi Zhou, Hui Zhang, Hana Lee, Shuyang Sun et al.CVPR 2022 · 20 citations
- Video Panoptic SegmentationDahun Kim, Sanghyun Woo, Joon-Young Lee, In So KweonCVPR 2020
- VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic SegmentationSiyuan Qiao, Yukun Zhu, Hartwig Adam, Alan L. Yuille et al.CVPR 2021
- PanoRecon: Real-Time Panoptic 3D Reconstruction from Monocular VideoDong Wu, Zike Yan, Hongbin ZhaCVPR 2024 · 8 citations
