LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
Rui Li, Biao Zhang, Zhenyu Li, Federico Tombari, Peter Wonka
Abstract
We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional depth estimation, which is limited to visible surfaces, LaRI predicts multiple surfaces intersected by the camera rays using layered point maps. Compared to the existing approaches that leverage neural implicit representations or iterative refinement, LaRI achieves complete scene reconstruction in one feed-forward pass, enabling efficient and view-aligned geometric reasoning to underpin both object-level and scene-level tasks. We further propose to predict the ray stopping index, which identifies valid intersecting pixels and layers from LaRI's output. To better underpin and evaluate this task, we build an annotation pipeline using rendering engines, construct annotations for five public datasets, including synthetic and real-world data covering 3D objects and scenes. As a generic method, LaRI's performance is validated in object-level and scene-level reconstruction tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f957b43-696f-46f5-8f7e-3c01ee1494beCited by top-tier papers3
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 17 citations
- RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object CompletionBardienus Pieter Duisterhof, Jan Oberst, Bowen Wen, Stan Birchfield et al.NeurIPS 2025 · 8 citations
- Pixal3D: Pixel-Aligned 3D Generation from ImagesDong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu et al.SIGGRAPH 2026 · 1 citation
Builds on43
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Object-Driven Multi-Layer Scene Decomposition From a Single ImageHelisa Dhamo, Nassir Navab, Federico TombariICCV 2019 · 40 citations
- Virtual Occlusions Through Implicit DepthJamie Watson, Mohamed Sayed, Zawar Qureshi, Gabriel J. Brostow et al.CVPR 2023
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng et al.CVPR 2021
- Footprints and Free Space From a Single Color ImageJamie Watson, Michael Firman, Áron Monszpart, Gabriel J. BrostowCVPR 2020
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
