RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
Bardienus Pieter Duisterhof, Jan Oberst, Bowen Wen, Stan Birchfield, Deva Ramanan, Jeffrey Ichnowski
Abstract
3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sharp object boundaries. Our work (RaySt3R) addresses these limitations by recasting 3D shape completion as a novel view synthesis problem. Specifically, given a single RGB-D image and a novel viewpoint (encoded as a collection of query rays), we train a feedforward transformer to predict depth maps, object masks, and per-pixel confidence scores for those query rays. RaySt3R fuses these predictions across multiple query views to reconstruct complete 3D shapes. We evaluate RaySt3R on synthetic and real-world datasets, and observe it achieves state-of-the-art performance, outperforming the baselines on all datasets by up to 44% in 3D chamfer distance. Project page: https://rayst3r.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a78c01e-afbf-4cc9-9b63-7d5be7204926Cited by top-tier papers6
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu et al.CVPR 2026 · 87 citations
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- Points-to-3D: Structure-Aware 3D Generation with Point Cloud PriorsJiatong Xia, Zicheng Duan, Anton van den Hengel, Lingqiao LiuCVPR 2026 · 6 citations
- Pixal3D: Pixel-Aligned 3D Generation from ImagesDong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu et al.SIGGRAPH 2026 · 1 citation
- LaRI: Layered Ray Intersections for Single-view 3D Geometric ReasoningRui Li, Biao Zhang, Zhenyu Li, Federico Tombari et al.ICML 2026
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- 3D Scene Reconstruction With Multi-Layer Depth and Epipolar TransformersDaeyun Shin, Zhile Ren, Erik B. Sudderth, Charless C. FowlkesICCV 2019 · 67 citations
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesXiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam et al.CVPR 2026 · 28 citations
- Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit ReconstructionLuoxi Zhang, Pragyan Shrestha, Yu Zhou, Chun Xie et al.ICCV 2025
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
- ART: Articulated Reconstruction TransformerZizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins et al.CVPR 2026 · 12 citations
