ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, Jiajun Wu
Abstract
We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose “SDS anchoring” to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Code and models are available at this url.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5bfa2e3-c5ed-4117-8aaf-3dc1c2eaab8eCited by top-tier papers52
- CAT3D: Create Anything in 3D with Multi-View Diffusion ModelsRuiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee et al.NeurIPS 2024 · 490 citations
- World-In-World: World Models in a Closed-Loop WorldJiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu et al.ICLR 2026 · 46 citations
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion PriorsXingyilang Yin, Qi Zhang, Jiahao Chang, Ying Feng et al.ICML 2026 · 33 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
Builds on27
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- Nerfbusters: Removing Ghostly Artifacts from Casually Captured NeRFsFrederik Warburg, Ethan Weber, Matthew Tancik, Aleksander Holynski et al.ICCV 2023 · 92 citations
- CAD : Photorealistic 3D Generation via Adversarial DistillationZiyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu et al.CVPR 2024 · 3 citations
- Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsZixin Zou, Weihao Cheng, Yan-Pei Cao, Shi-Sheng Huang et al.AAAI 2024 · 34 citations
- MVIP-NeRF: Multi-View 3D Inpainting on NeRF Scenes via Diffusion PriorHonghua Chen, Chen Change Loy, Xingang PanCVPR 2024
- How to Use Diffusion Priors under Sparse Views?Qisen Wang, Yifan Zhao, Jiawei Ma, Jia LiNeurIPS 2024 · 12 citations
