Multi-view Consistent 3D Panoptic Scene Understanding
Xianzhu Liu, Xin Sun, Haozhe Xie, Zonglin Li, Ru Li, Shengping Zhang
Abstract
3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machinegenerated 2D panoptic segmentation masks as training labels. However, these generated masks often exhibit multi-view inconsistencies, leading to ambiguities during the optimization process. To address this, we present Multi-view Consistent 3D Panoptic Scene Understanding (MVC-PSU), featuring two key components: 1) Probabilistic Semantic Aligner, which associates semantic information of corresponding pixels across multiple views by probabilistic alignment to ensure that the predicted panoptic segmentation masks are consistent across different views. 2) Geometric Consistency Enforcer, which uses multi-view projection and monocular depth consistency to ensure that the geometry of the reconstructed scene is accurate and consistent across different views. Experimental results demonstrate that the proposed MVC-PSU surpasses state-of-the-art methods on the ScanNet, Replica, and HyperSim datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5822d477-586e-46db-be27-cb831f864e52Cited by top-tier papers2
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object CorrespondenceJiancheng Pan, Runze Wang, Tianwen Qian, Mohammad Mahdi et al.CVPR 2026
- Evolving Semantic Propagation for Aerial Semantic 3D Gaussian SplattingZihan Gao, Lingling Li, Xu Liu, Fang Liu et al.AAAI 2026
Builds on31
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth-supervised NeRF: Fewer Views and Faster Training for FreeKangle Deng, Andrew Liu, Jun-Yan Zhu, Deva RamananCVPR 2022 · 756 citations
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 551 citations
Related papers
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu et al.ICML 2026 · 3 citations
- PanSt3R: Multi-View Consistent Panoptic SegmentationLojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld et al.ICCV 2025 · 5 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesXiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam et al.CVPR 2026 · 28 citations
- BUOL: A Bottom-Up Framework with Occupancy-Aware Lifting for Panoptic 3D Scene Reconstruction From a Single ImageTao Chu, Pan Zhang, Qiong Liu, Jiaqi WangCVPR 2023
