DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
Qitao Zhao, Amy Lin, Jeff Tan, Jason Y. Zhang, Deva Ramanan, Shubham Tulsiani
Abstract
Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning approach that directly infers 3D scene geometry and camera poses from multi-view images. Our framework, Dif-fusionSfM, parameterizes scene geometry and cameras as pixel-wise ray origins and endpoints in a global frame and employs a transformer-based denoising diffusion model to predict them from multi-view inputs. To address practical challenges in training diffusion models with missing data and unbounded scene coordinates, we introduce specialized mechanisms that ensure robust learning. We empirically validate DiffusionSfM on both synthetic and real datasets, demonstrating that it outperforms classical and learningbased approaches while naturally modeling uncertainty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b54eb24e-3ea1-4afd-9638-fbe50fd5648cCited by top-tier papers4
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- Flow3r: Factored Flow Prediction for Scalable Visual Geometry LearningZhongxiao Cong, Qitao Zhao, Minsik Jeon, Shubham TulsianiCVPR 2026 · 8 citations
- MonoFusion: Sparse-View 4D Reconstruction via Monocular FusionZihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri et al.ICCV 2025 · 5 citations
- Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature AlignmentYouming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal et al.CVPR 2026 · 4 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
Related papers
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang et al.ICLR 2024 · 126 citations
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- DiffSF: Diffusion Models for Scene Flow EstimationYushan Zhang, Bastian Wandt, Maria Magnusson, Michael FelsbergNeurIPS 2024 · 8 citations
- RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and GenerationTitas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson et al.CVPR 2023
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose EstimationJunwen Huang, Shishir Reddy Vutukur, Peter KT Yu, Nassir Navab et al.ICCV 2025 · 1 citation
