GNeSF: Generalizable Neural Semantic Fields
Hanlin Chen, Chen Li, Mengqi Guo, Zhiwen Yan, Gim Hee Lee
Abstract
3D scene segmentation based on neural implicit representation has emerged recently with the advantage of training only on 2D supervision. However, existing approaches still requires expensive per-scene optimization that prohibits generalization to novel scenes during inference. To circumvent this problem, we introduce a generalizable 3D segmentation framework based on implicit representation. Specifically, our framework takes in multi-view image features and semantic maps as the inputs instead of only spatial information to avoid overfitting to scene-specific geometric and semantic information. We propose a novel soft voting mechanism to aggregate the 2D semantic information from different views for each 3D point. In addition to the image features, view difference information is also encoded in our framework to predict the voting scores. Intuitively, this allows the semantic information from nearby views to contribute more compared to distant ones. Furthermore, a visibility module is also designed to detect and filter out detrimental information from occluded views. Due to the generalizability of our proposed method, we can synthesize semantic maps or conduct 3D semantic segmentation for novel scenes with solely 2D semantic supervision. Experimental results show that our approach achieves comparable performance with scene-specific approaches. More importantly, our approach can even outperform existing strong supervision-based approaches with only 2D annotations. Our source code is available at: https://github.com/HLinChen/GNeSF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ad18cfd-42fd-41fc-a813-1fbdc76485deCited by top-tier papers4
- VCR-GauS: View Consistent Depth-Normal Regularizer for Gaussian Surface ReconstructionHanlin Chen, Fangyin Wei, Chen Li, Tianxin Huang et al.NeurIPS 2024 · 71 citations
- PE3R: Perception-Efficient 3D ReconstructionJie Hu, Shizun Wang, Xinchao WangCVPR 2026 · 9 citations
- WildSeg3D: Segment Any 3D Objects in the Wild from 2D ImagesYansong Guo, Jie Hu, Yansong Qu, Liujuan CaoICCV 2025 · 1 citation
- SANeRF-HQ: Segment Anything for NeRF in High QualityYichen Liu, Benran Hu, Chi-Keung Tang, Yu-Wing TaiCVPR 2024
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Fully Convolutional Geometric FeaturesChristopher B. Choy, Jaesik Park, Vladlen KoltunICCV 2019 · 807 citations
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 551 citations
Related papers
- GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic FieldsYunsong Wang, Hanlin Chen, Gim Hee LeeCVPR 2024 · 2 citations
- GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene UnderstandingZi-Ting Chou, Sheng-Yu Huang, I-Jieh Liu, Yu-Chiang Frank WangCVPR 2024
- GS2-GNeSF: Geometry-Semantics Synergy for Generalizable Neural Semantic FieldsChengshun Wang, Na ZhaoACM MM 2024
- GRF: Learning a General Radiance Field for 3D Representation and RenderingAlex Trevithick, Bo YangICCV 2021 · 258 citations
- OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive LearningHaiyang Ying, Yixuan Yin, Jinzhi Zhang, Fan Wang et al.CVPR 2024 · 32 citations
