MultiDiff: Consistent Novel View Synthesis from a Single Image
Norman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi, Samuel Rota Bulò, Matthias Nießner, Peter Kontschieder
Abstract
We introduce MultiDiff, a novel approach for consistent novel view synthesis of scenes from a single RGB image. The task of synthesizing novel views from a single reference image is highly ill-posed by nature, as there exist multiple, plausible explanations for unobserved areas. To address this issue, we incorporate strong priors in form of monocular depth predictors and video-diffusion models. Monocular depth enables us to condition our model on warped reference images for the target views, increasing geometric stability. The video-diffusion prior provides a strong proxy for 3D scenes, allowing the model to learn continuous and pixel-accurate correspondences across generated images. In contrast to approaches relying on autoregressive image generation that are prone to drifts and error accumulation, MultiDiff Jointly synthesizes a sequence of frames yielding high-quality and multi-view consistent results - even for long-term scene generation with large camera movements, while reducing inference time by an order of magnitude. For additional consistency and image quality improvements, we introduce a novel, structured noise distribution. Our experimental results demonstrate that MultiDiff outperforms state-of-the-art methods on the challenging, real-world datasets RealEstate10K and ScanNet. Finally, our model naturally supports multi-view consistent editing without the need for further tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers40
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsMark Yu, Wenbo Hu, Jinbo Xing, Ying ShanICCV 2025 · 25 citations
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video GuidanceZun Wang, Jaemin Cho, Jialu Li, Han Lin et al.ICML 2026 · 19 citations
- Vista4D: Video Reshooting with 4D Point CloudsKuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant et al.CVPR 2026 · 17 citations
- Dynamic View Synthesis as an Inverse ProblemHidir Yesiltepe, Pinar YanardagNeurIPS 2025 · 12 citations
- Bolt3D: Generating 3D Scenes in SecondsStanislaw Szymanowicz, Jason Y. Zhang, Pratul P. Srinivasan, Ruiqi Gao et al.ICCV 2025 · 11 citations
Builds on54
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionXueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh KhoshelhamACM MM 2025
- Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesSarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang et al.ICCV 2025 · 3 citations
- Generative Novel View Synthesis with 3D-Aware Diffusion ModelsEric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman et al.ICCV 2023 · 314 citations
- Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionHaiping Wang, Yuan Liu, Ziwei Liu, Wenping Wang et al.ICCV 2025 · 8 citations
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsSongchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie et al.ICCV 2025 · 6 citations
