VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic Segmentation
Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen
Abstract
In this paper, we present ViP-DeepLab, a unified model attempting to tackle the long-standing and challenging inverse projection problem in vision, which we model as restoring the point clouds from perspective image sequences while providing each point with instance-level semantic interpretations. Solving this problem requires the vision models to predict the spatial location, semantic class, and temporally consistent instance label for each 3D point. ViP-DeepLab approaches it by jointly performing monocular depth estimation and video panoptic segmentation. We name this joint task as Depth-aware Video Panoptic Segmentation, and propose a new evaluation metric along with two derived datasets for it, which will be made available to the public. On the individual sub-tasks, ViP-DeepLab also achieves state-of-the-art results, outperforming previous methods by 5.1% VPQ on Cityscapes-VPS, ranking 1st on the KITTI monocular depth estimation benchmark, and 1st on KITTI MOTS pedestrian. The datasets and the evaluation codes are made publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2023 · 285 citations
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- Tracking Anything with Decoupled Video SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing et al.ICCV 2023 · 240 citations
- IEBins: Iterative Elastic Bins for Monocular Depth EstimationShuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu et al.NeurIPS 2023 · 114 citations
Builds on17
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- Spatial-Temporal Relation Networks for Multi-Object TrackingJiarui Xu, Yue Cao, Zheng Zhang, Han HuICCV 2019 · 260 citations
Related papers
- MGNet: Monocular Geometric Scene Understanding for Autonomous DrivingMarkus Schön, Michael Buchholz, Klaus DietmayerICCV 2021 · 60 citations
- Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance LearningJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu et al.ICCV 2023 · 11 citations
- Video Panoptic SegmentationDahun Kim, Sanghyun Woo, Joon-Young Lee, In So KweonCVPR 2020
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li et al.CVPR 2022 · 58 citations
- PanopticDepth: A Unified Framework for Depth-aware Panoptic SegmentationNaiyu Gao, Fei He, Jian Jia, Yanhu Shan et al.CVPR 2022 · 27 citations
