DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
Letian Wang, Seung Wook Kim, Jiawei Yang, Cunjun Yu, Boris Ivanovic, Steven L. Waslander, Yue Wang, Sanja Fidler, Marco Pavone, Péter Karkus
Abstract
We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, single-frame multi-view camera inputs with limited view overlap, and is trained self-supervised with differentiable rendering to reconstruct RGB, depth, or feature images. Our first insight is to exploit per-scene optimized Neural Radiance Fields (NeRFs) by generating dense depth and virtual camera targets from them, which helps our model to learn enhanced 3D geometry from sparse non-overlapping image inputs. Second, to learn a semantically rich 3D representation, we propose distilling features from pre-trained 2D foundation models, such as CLIP or DINOv2, thereby enabling various downstream tasks without the need for costly 3D human annotations. To leverage these two insights, we introduce a novel model architecture with a two-stage lift-splat-shoot encoder and a parameterized sparse hierarchical voxel representation. Experimental results on the NuScenes and Waymo NOTR datasets demonstrate that DistillNeRF significantly outperforms existing comparable state-of-the-art self-supervised methods for scene reconstruction, novel view synthesis, and depth estimation; and it allows for competitive zero-shot 3D semantic occupancy prediction, as well as open-world scene understanding through distilled foundation model features. Demos and code will be available at https://distillnerf.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c0e4efe-76c2-4a56-b9c0-d5ec90a47cbbCited by top-tier papers16
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingYang Zhou, Hao Shao, Letian Wang, Zhuofan Zong et al.ICLR 2026 · 20 citations
- See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy PredictionYuan Wu, Zhiqiang Yan, Yigong Zhang, Xiang Li et al.NeurIPS 2025 · 7 citations
- OccAny: Generalized Unconstrained Urban 3D OccupancyAnh-Quan Cao, Tuan-Hung VuCVPR 2026 · 6 citations
- Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy PredictionChi Yan, Dan XuICLR 2026 · 6 citations
- VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian SplattingHoonhee Cho, Jae-Young Kang, Giwon Lee, Hyemin Yang et al.NeurIPS 2025 · 5 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
Related papers
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu et al.NeurIPS 2023 · 173 citations
- FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation ModelsJianglong Ye, Naiyan Wang, Xiaolong WangICCV 2023 · 56 citations
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisAjay Jain, Matthew Tancik, Pieter AbbeelICCV 2021 · 615 citations
- To View Transform or Not to View Transform: NeRF-based Pre-training PerspectiveHyeonjun Jeong, Juyeb Shin, Dongsuk KumICLR 2026 · 2 citations
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 479 citations
