Panoptic Compositional Feature Field for Editable Scene Rendering with Network-Inferred Labels via Metric Learning
Xinhua Cheng, Yanmin Wu, Mengxi Jia, Qian Wang, Jian Zhang
Abstract
Despite neural implicit representations demonstrating impressive high-quality view synthesis capacity, decomposing such representations into objects for instance-level editing is still challenging. Recent works learn objectcompositional representations supervised by ground truth instance annotations and produce promising scene editing results. However, ground truth annotations are manually labeled and expensive in practice, which limits their usage in real-world scenes. In this work, we attempt to learn an object-compositional neural implicit representation for editable scene rendering by leveraging labels inferred from the off-the-shelf 2D panoptic segmentation networks instead of the ground truth annotations. We propose a novel framework named Panoptic Compositional Feature Field (PCFF), which introduces an instance quadruplet metric learning to build a discriminating panoptic feature space for reliable scene editing. In addition, we propose semanticrelated strategies to further exploit the correlations between semantic and appearance attributes for achieving better rendering results. Experiments on multiple scene datasets including ScanNet, Replica, and ToyDesk demonstrate that our proposed method achieves superior performance for novel view synthesis and produces convincing real-world scene editing results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li et al.ICLR 2024 · 58 citations
- OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive LearningHaiyang Ying, Yixuan Yin, Jinzhi Zhang, Fan Wang et al.CVPR 2024 · 32 citations
- Aerial Lifting: Neural Urban Semantic and Building Instance Lifting from Aerial ImageryYuqi Zhang, Guanying Chen, Jiaxing Chen, Shuguang CuiCVPR 2024 · 4 citations
- Retri3D: 3D Neural Graphics Representation RetrievalYushi Guan, Daniel Kwan, Jean Sebastien Dandurand, Xi Yan et al.ICLR 2025
Builds on30
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li et al.ICCV 2021 · 1,284 citations
Related papers
- Panoptic Neural Fields: A Semantic Object-Aware Neural Scene RepresentationAbhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi et al.CVPR 2022 · 204 citations
- Learning Object-Compositional Neural Radiance Field for Editable Scene RenderingBangbang Yang, Yinda Zhang, Yinghao Xu, Yijin Li et al.ICCV 2021 · 305 citations
- AssetField: Assets Mining and Reconfiguration in Ground Feature Plane RepresentationYuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao et al.ICCV 2023 · 13 citations
- Learning Unified Decompositional and Compositional NeRF for Editable Novel View SynthesisYuxin Wang, Wayne Wu, Dan XuICCV 2023 · 18 citations
- RICO: Regularizing the Unobservable for Indoor Compositional ReconstructionZizhang Li, Xiaoyang Lyu, Yuanyuan Ding, Mengmeng Wang et al.ICCV 2023 · 17 citations
