SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction
Yuanhui Huang, Wenzhao Zheng, Borui Zhang, Jie Zhou, Jiwen Lu
Abstract
Only Video Sequences As Supervision Self-Supervised 3D Occupancy Prediction Semantics Geometry lower higher …… SelfOcc Optional 2D Segmentor car driveable surface sidewalk terrain manmade vegetation truck Figure 1 . Trained with only video sequences as supervision, our model can predict meaningful geometry for the scene given surroundcamera RGB images, which can be further extended to semantic occupancy prediction if 2D segmentation maps are available e.g. from an off-the-shelf segmentor. This task is challenging because it completely depends on video sequences to reconstruct scenes without any 3D supervision. We observe that our model can produce dense and consistent occupancy prediction and even infer the back side of cars.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26314514-32eb-48be-9d53-6e9b50944de4Cited by top-tier papers49
- OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree QueriesYuhang Lu, Xinge Zhu, Tai Wang, Yuexin MaNeurIPS 2024 · 70 citations
- RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging RadarFangqiang Ding, Xiangyu Wen, Yunzhou Zhu, Yiming Li et al.NeurIPS 2024 · 66 citations
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- RAP: 3D Rasterization Augmented End-to-End PlanningLan Feng, Yang Gao, Eloi Zablocki, Quanyi Li et al.ICLR 2026 · 47 citations
- DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view InputQijian Tian, Xin Tan, Yuan Xie, Lizhuang MaAAAI 2025 · 45 citations
Builds on39
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy EstimationSimon Boeder, Fabian Gigengack, Simon Roesler, Holger Caesar et al.CVPR 2026 · 7 citations
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model GuidanceDuc-Hai Pham, Duc Dung Nguyen, Anh Pham, Tuan Ho et al.AAAI 2025 · 6 citations
- OccAny: Generalized Unconstrained Urban 3D OccupancyAnh-Quan Cao, Tuan-Hung VuCVPR 2026 · 6 citations
- QueryOcc: Query-based Self-Supervision for 3D Semantic OccupancyAdam Lilja, Ji Lan, Junsheng Fu, Lars HammarstrandCVPR 2026 · 4 citations
