Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
Aleksandar Jevtic, Christoph Reich, Felix Wimbauer, Oliver Hahn, Christian Rupprecht, Stefan Roth, Daniel Cremers
Abstract
Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from self-supervised representation learning and 2D unsupervised scene understanding to SSC. Our training exclusively utilizes multi-view consistency self-supervision without any form of semantic or geometric ground truth. Given a single input image, SceneDINO infers the 3D geometry and expressive 3D DINO features in a feed-forward manner. Through a novel 3D feature distillation approach, we obtain unsupervised 3D semantics. In both 3D and 2D unsupervised scene understanding, SceneDINO reaches state-of-the-art segmentation accuracy. Linear probing our 3D features matches the segmentation accuracy of a current supervised SSC approach. Additionally, we showcase the domain generalization and multi-view consistency of SceneDINO, taking the first steps towards a strong foundation for single image 3D scene understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa75ccdb-37af-4cd6-8e45-0ec69fd785baCited by top-tier papers6
- DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed ImagesXiaoxue Chen, Ziyi Xiong, Yuantao Chen, Gen Li et al.CVPR 2026 · 24 citations
- INSID3: Training-Free In-Context Segmentation with DINOv3Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers et al.CVPR 2026 · 13 citations
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic CameraHao Shi, Ze Wang, Shangwei Guo, Mengfei Duan et al.CVPR 2026 · 11 citations
- OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial PerspectiveMarkus Gross, Sai B. Matha, Aya Fahmy, Rui Song et al.CVPR 2026 · 7 citations
- OccAny: Generalized Unconstrained Urban 3D OccupancyAnh-Quan Cao, Tuan-Hung VuCVPR 2026 · 6 citations
Builds on44
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- UnScene3D: Unsupervised 3D Instance Segmentation for Indoor ScenesDávid Rozenberszki, Or Litany, Angela DaiCVPR 2024 · 25 citations
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
- Self-Supervised Image Representation Learning with Geometric Set ConsistencyNenglun Chen, Lei Chu, Hao Pan, Yan Lu et al.CVPR 2022 · 8 citations
- Multiview Compressive Coding for 3D ReconstructionChao-Yuan Wu, Justin Johnson, Jitendra Malik, Christoph Feichtenhofer et al.CVPR 2023
