Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
Rui Li, Tobias Fischer, Mattia Segù, Marc Pollefeys, Luc Van Gool, Federico Tombari
Abstract
Recovering the 3D scene geometry from a single view is a fundamental yet ill-posed problem in computer vision. While classical depth estimation methods infer only a 2.5D scene representation limited to the image plane, recent approaches based on radiance fields reconstruct a full 3D representation. However, these methods still struggle with occluded regions since inferring geometry without visual observation requires (i) semantic knowledge of the surroundings, and (ii) reasoning about spatial context. We propose KYN, a novel method for single-view scene reconstruction that reasons about semantic and spatial context to predict each point's density. We introduce a vision-language modulation module to enrich point features with fine-grained semantic information. We aggregate point representations across the scene through a language-guided spatial attention mechanism to yield per-point density predictions aware of the 3D semantic context. We show that KYN improves 3D shape recovery compared to predicting density for each 3D point in isolation. We achieve state-of-the-art results in scene and object reconstruction on KITTI-360, and show improved zero-shot generalization compared to prior work. Project page: https://ruili3.github.io/kyn .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7618dfc-265a-463d-b2e2-7d12e9370a6aCited by top-tier papers8
- LRM-Zero: Training Large Reconstruction Models with Synthesized DataDesai Xie, Sai Bi, Zhixin Shu, Kai Zhang et al.NeurIPS 2024 · 36 citations
- ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionYi Feng, Yu Han, Xijing Zhang, Tanghui Li et al.AAAI 2025 · 8 citations
- Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesSarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang et al.ICCV 2025 · 3 citations
- The Midas Touch for Metric DepthYu Ma, Zizhan Guo, Zuyi Xiong, Haoran Zhang et al.CVPR 2026 · 2 citations
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
Builds on32
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View ImageWonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee et al.ICCV 2025 · 2 citations
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- Single Image 3D Object Estimation with Primitive Graph NetworksQian He, Desen Zhou, Bo Wan, Xuming HeACM MM 2021 · 1 citation
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin et al.AAAI 2025 · 11 citations
