NeRF-SOS: Any-View Self-supervised Object Segmentation on Complex Scenes
Zhiwen Fan, Peihao Wang, Yifan Jiang, Xinyu Gong, Dejia Xu, Zhangyang Wang
Abstract
Neural volumetric representations have shown the potential that Multi-layer Perceptrons (MLPs) can be optimized with multi-view calibrated images to represent scene geometry and appearance without explicit 3D supervision. Object segmentation can enrich many downstream applications based on the learned radiance field. However, introducing hand-crafted segmentation to define regions of interest in a complex real-world scene is non-trivial and expensive as it acquires per view annotation. This paper carries out the exploration of self-supervised learning for object segmentation using NeRF for complex real-world scenes. Our framework, called NeRF with Self-supervised Object Segmentation (NeRF-SOS), couples object segmentation and neural radiance field to segment objects in any view within a scene. By proposing a novel collaborative contrastive loss in both appearance and geometry levels, NeRF-SOS encourages NeRF models to distill compact geometry-aware segmentation clusters from their density fields and the self-supervised pre-trained 2D visual features. The self-supervised object segmentation framework can be applied to various NeRF models that both lead to photo-realistic rendering results and convincing segmentation maps for both indoor and outdoor scenarios. Extensive results on the LLFF, BlendedMVS, CO3Dv2, and Tank & Temples datasets validate the effectiveness of NeRF-SOS. It consistently surpasses other 2D-based self-supervised baselines and predicts finer object masks than existing supervised counterparts. Please refer to the video on our project page for more details: https://zhiwenfan.github.io/NeRF-SOS/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c960a7c4-8a1a-49b9-91c1-273d6e5ce89bCited by top-tier papers37
- LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPSZhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu et al.NeurIPS 2024 · 681 citations
- Segment Anything in 3D with NeRFsJiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang et al.NeurIPS 2023 · 255 citations
- Segment Any 3D GaussiansJiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie et al.AAAI 2025 · 175 citations
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu et al.NeurIPS 2023 · 173 citations
- Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature FieldsShijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan et al.CVPR 2024 · 145 citations
Builds on34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
Related papers
- Unsupervised Multi-View Object Segmentation Using Radiance Field PropagationXinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai et al.NeurIPS 2022 · 34 citations
- Instance Neural Radiance FieldYichen Liu, Benran Hu, Junkai Huang, Yu-Wing Tai et al.ICCV 2023 · 49 citations
- SANeRF-HQ: Segment Anything for NeRF in High QualityYichen Liu, Benran Hu, Chi-Keung Tang, Yu-Wing TaiCVPR 2024
- ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldZhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang et al.AAAI 2024 · 18 citations
- SUDS: Scalable Urban Dynamic ScenesHaithem Turki, Jason Y. Zhang, Francesco Ferroni, Deva RamananCVPR 2023
