Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views
Zixin Zou, Weihao Cheng, Yan-Pei Cao, Shi-Sheng Huang, Ying Shan, Song-Hai Zhang
Abstract
Reconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distilling pre-trained diffusion priors into 3D representations using score distillation sampling (SDS), these methods often struggle to simultaneously achieve high-quality, consistent, and detailed results for both novel-view synthesis (NVS) and geometry. In this work, we present Sparse3D, a novel 3D reconstruction method tailored for sparse view inputs. Our approach distills robust priors from a multiview-consistent diffusion model to refine a neural radiance field. Specifically, we employ a controller that harnesses epipolar features from input views, guiding a pre-trained diffusion model, such as Stable Diffusion, to produce novel-view images that maintain 3D consistency with the input. By tapping into 2D priors from powerful image diffusion models, our integrated model consistently delivers high-quality results, even when faced with open-world objects. To address the blurriness introduced by conventional SDS, we introduce the category-score distillation sampling (C-SDS) to enhance detail. We conduct experiments on CO3DV2 which is a multi-view dataset of real-world objects. Both quantitative and qualitative evaluations demonstrate that our approach outperforms previous state-of-the-art works on the metrics regarding NVS and geometry reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a9d38ce-416f-4cc6-a951-f9bb3d21793bCited by top-tier papers12
- Director3D: Real-world Camera Trajectory and 3D Scene Generation from TextXinyang Li, Zhangyu Lai, Linning Xu, Yansong Qu et al.NeurIPS 2024 · 60 citations
- Rethinking Score Distillation as a Bridge Between Image DistributionsDavid McAllister, Songwei Ge, Jia-Bin Huang, David Jacobs et al.NeurIPS 2024 · 43 citations
- ConsistNet: Enforcing 3D Consistency for Multi-View Images DiffusionJiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji et al.CVPR 2024 · 25 citations
- Denoising Diffusion via Image-Based RenderingTitas Anciukevicius, Fabian Manhardt, Federico Tombari, Paul HendersonICLR 2024 · 19 citations
- HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3DSangmin Woo, Byeongjun Park, Hyojun Go, Jin-Young Kim et al.CVPR 2024 · 11 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- How to Use Diffusion Priors under Sparse Views?Qisen Wang, Yifan Zhao, Jiawei Ma, Jia LiNeurIPS 2024 · 12 citations
- CAD : Photorealistic 3D Generation via Adversarial DistillationZiyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu et al.CVPR 2024 · 3 citations
- Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationJunyoung Seo, Wooseok Jang, Minseop Kwak, Inès Hyeonsu Kim et al.ICLR 2024 · 157 citations
- MVIP-NeRF: Multi-View 3D Inpainting on NeRF Scenes via Diffusion PriorHonghua Chen, Chen Change Loy, Xingang PanCVPR 2024
- Wonder3D: Single Image to 3D Using Cross-Domain DiffusionXiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu et al.CVPR 2024 · 269 citations
