ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
Simon Boeder, Fabian Gigengack, Simon Roesler, Holger Caesar, Benjamin Risse
摘要
Recent progress in self- and weakly supervised occupancy estimation has largely relied on 2D projection or rendering-based supervision, which suffers from geometric inconsistencies and severe depth bleeding. We thus introduce ShelfOcc, a vision-only method that overcomes these limitations without relying on LiDAR. ShelfOcc brings supervision into native 3D space by generating metrically consistent semantic voxel labels from video, enabling true 3D supervision without any additional sensors or manual 3D annotations. While recent vision-based 3D geometry foundation models provide a promising source of prior knowledge, they do not work out of the box as a prediction due to sparse or noisy and inconsistent geometry, especially in dynamic driving scenes. Our method introduces a dedicated framework that mitigates these issues by filtering and accumulating static geometry consistently across frames, handling dynamic content and propagating semantic information into a stable voxel representation. This data-centric shift in supervision for weakly/shelf-supervised occupancy estimation allows the use of essentially any SOTA occupancy model architecture without relying on LiDAR data. We argue that such high-quality supervision is essential for robust occupancy learning and constitutes an important complementary avenue to architectural innovation. On the Occ3D-nuScenes benchmark, ShelfOcc substantially outperforms all previous weakly/shelf-supervised methods (up to a 34% relative improvement), establishing a new data-driven direction for LiDAR-free 3D scene understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie 等CVPR 2024 · 被引用 513 次
相关 Paper
- Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model GuidanceDuc-Hai Pham, Duc Dung Nguyen, Anh Pham, Tuan Ho 等AAAI 2025 · 被引用 6 次
- QueryOcc: Query-based Self-Supervision for 3D Semantic OccupancyAdam Lilja, Ji Lan, Junsheng Fu, Lars HammarstrandCVPR 2026 · 被引用 4 次
- SelfOcc: Self-Supervised Vision-Based 3D Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Borui Zhang, Jie Zhou 等CVPR 2024
- GS-Occ3D: Scaling Vision-Only Occupancy Reconstruction with Gaussian SplattingBaijun Ye, Minghui Qin, Saining Zhang, Moonjun Goon 等ICCV 2025 · 被引用 2 次
- Test-Time 3D Occupancy PredictionFengyi Zhang, Xiangyu Sun, Huitong Yang, Zheng Zhang 等CVPR 2026 · 被引用 2 次
