Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
Vitor Guizilini, Muhammad Zubair Irshad, Dian Chen, Greg Shakhnarovich, Rares Ambrus
Abstract
Figure 1 . MVGD is a state-of-the-art method that generates images and scale-consistent depth maps from novel viewpoints given an arbitrary number of posed input views. In the above, red cameras are used as conditioning to directly generate RGB-D predictions from green cameras. To highlight the multi-view consistency of our method, predicted colored pointclouds from all novel viewpoints are stacked together for visualization without any post-processing. More examples and videos can be found in https://mvgd.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cba3bfe-d935-452a-8592-6fc04c9bb416Cited by top-tier papers4
- Cameras as Relative Positional EncodingRuilong Li, Brent Yi, Junchen Liu, Hang Gao et al.NeurIPS 2025 · 113 citations
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao et al.CVPR 2026 · 38 citations
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsMichal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang et al.NeurIPS 2025 · 5 citations
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image FeaturesAxel Barroso-Laguna, Tommaso Cavallari, Victor Prisacariu, Eric BrachmannICLR 2026 · 3 citations
Builds on61
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- MVD-Fusion: Single-view 3D via Depth-consistent Multi-view GenerationHanzhe Hu, Zhizhuo Zhou, Varun Jampani, Shubham TulsianiCVPR 2024
- MultiDiff: Consistent Novel View Synthesis from a Single ImageNorman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi et al.CVPR 2024 · 14 citations
- Self-Supervised Visibility Learning for Novel View SynthesisYujiao Shi, Hongdong Li, Xin YuCVPR 2021
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- Consistent View Synthesis with Pose-Guided Diffusion ModelsHung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan et al.CVPR 2023
