Fast Monocular Scene Reconstruction with Global-Sparse Local-Dense Grids
Wei Dong, Christopher B. Choy, Charles Loop, Or Litany, Yuke Zhu, Anima Anandkumar
Abstract
Indoor scene reconstruction from monocular images has long been sought after by augmented reality and robotics developers. Recent advances in neural field representations and monocular priors have led to remarkable results in scene-level surface reconstructions. The reliance on Multilayer Perceptrons (MLP), however, significantly limits speed in training and rendering. In this work, we propose to directly use signed distance function (SDF) in sparse voxel block grids for fast and accurate scene reconstruction without MLPs. Our globally sparse and locally dense data structure exploits surfaces' spatial sparsity, enables cache-friendly queries, and allows direct extensions to multi-modal data such as color and semantic labels. To apply this representation to monocular scene reconstruction, we develop a scale calibration algorithm for fast geometric initialization from monocular depth priors. We apply differentiable volume rendering from this initialization to refine details with fast convergence. We also introduce efficient high-dimensional Continuous Random Fields (CRFs) to further exploit the semantic-geometry consistency between scene objects. Experiments show that our approach is 10× faster in training and 100× faster in rendering while achieving comparable accuracy to state-of-the-art neural implicit methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b663e5a-cfb0-4418-aaa2-bd68ae21bd69Cited by top-tier papers2
- Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionGuoyu LuCVPR 2025
- NC-SDF: Enhancing Indoor Scene Reconstruction Using Neural SDFs with View-Dependent Normal CompensationZiyi Chen, Xiaolong Wu, Yu ZhangCVPR 2024
Builds on24
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
Related papers
- Neural Geometric Level of Detail: Real-Time Rendering With Implicit 3D ShapesTowaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis et al.CVPR 2021
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view ReconstructionQiancheng Fu, Qingshan Xu, Yew Soon Ong, Wenbing TaoNeurIPS 2022 · 336 citations
- NeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene ReconstructionWenyuan Zhang, Emily Yue-ting Jia, Junsheng Zhou, Baorui Ma et al.CVPR 2025
- VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View ReconstructionYufan Ren, Fangjinhua Wang, Tong Zhang, Marc Pollefeys et al.CVPR 2023
