Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
Duc-Hai Pham, Duc Dung Nguyen, Anh Pham, Tuan Ho, Phong Nguyen, Khoi Nguyen, Rang Nguyen
Abstract
Accurate prediction of 3D semantic occupancy from 2D visual images is crucial for enabling autonomous agents to understand their surroundings for planning and navigation. State-of-the-art methods typically rely on fully supervised approaches, requiring large labeled datasets obtained through expensive LiDAR sensors and meticulous voxel-wise annotation by human experts. The resource-intensive nature of this annotation process significantly limits the scalability and application of these methods. To address this challenge, we propose a novel semi-supervised framework that reduces reliance on densely annotated data. Our approach leverages 2D foundation models to extract essential 3D scene geometry and semantic cues, enabling a more efficient training process. The proposed framework has two key advantages: (1) Generalizability, as it is compatible with various 3D semantic scene completion methods, including 2D-3D lifting and 3D-2D transformer techniques; and (2) Effectiveness, as demonstrated by experiments on the SemanticKITTI and NYUv2 datasets, where our method achieves up to 85% of the fully supervised performance using only 10% of the labeled data. This approach not only reduces the cost of data annotation but also highlights its potential for broader adoption in visionbased systems for 3D semantic occupancy prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68d3ac88-ffae-4d79-830f-086b5e537bc4Builds on27
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy EstimationSimon Boeder, Fabian Gigengack, Simon Roesler, Holger Caesar et al.CVPR 2026 · 7 citations
- SelfOcc: Self-Supervised Vision-Based 3D Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Borui Zhang, Jie Zhou et al.CVPR 2024
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- Test-Time 3D Occupancy PredictionFengyi Zhang, Xiangyu Sun, Huitong Yang, Zheng Zhang et al.CVPR 2026 · 2 citations
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao et al.CVPR 2023
