How to Use Diffusion Priors under Sparse Views?
Qisen Wang, Yifan Zhao, Jiawei Ma, Jia Li
Abstract
Novel view synthesis under sparse views has been a long-term important challenge in 3D reconstruction. Existing works mainly rely on introducing external semantic or depth priors to supervise the optimization of 3D representations. However, the diffusion model, as an external prior that can directly provide visual supervision, has always underperformed in sparse-view 3D reconstruction using Score Distillation Sampling (SDS) due to the low information entropy of sparse views compared to text, leading to optimization challenges caused by mode deviation. To this end, we present a thorough analysis of SDS from the mode-seeking perspective and propose Inline Prior Guided Score Matching (IPSM), which leverages visual inline priors provided by pose relationships between viewpoints to rectify the rendered image distribution and decomposes the original optimization objective of SDS, thereby offering effective diffusion visual guidance without any fine-tuning or pre-training. Furthermore, we propose the IPSM-Gaussian pipeline, which adopts 3D Gaussian Splatting as the backbone and supplements depth and geometry consistency regularization based on IPSM to further improve inline priors and rectified distribution. Experimental results on different public datasets show that our method achieves state-of-the-art reconstruction quality. The code is released at https://github.com/iCVTEAM/IPSM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ee78905-edda-4c51-85ba-6d311e89c91eCited by top-tier papers2
- Novel View Synthesis from A Few Glimpses via Test-Time Natural Video CompletionYan Xu, Yixing Wang, Stella X. YuNeurIPS 2025 · 4 citations
- SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse RenderingJiahao Niu, Rongjia Zheng, Wenju Xu, Wei-Shi Zheng et al.CVPR 2026 · 1 citation
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsZixin Zou, Weihao Cheng, Yan-Pei Cao, Shi-Sheng Huang et al.AAAI 2024 · 34 citations
- CAD : Photorealistic 3D Generation via Adversarial DistillationZiyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu et al.CVPR 2024 · 3 citations
- GeoQuery: Geometry-Query Diffusion for Sparse-View ReconstructionXiao Cao, Yuze Li, Youmin Zhang, Jiayu Song et al.SIGGRAPH 2026
- ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent SynthesisXiangjun Gao, Xiaoyu Li, Chaopeng Zhang, Qi Zhang et al.CVPR 2024
- PR-IQA: Partial-Reference Image Quality Assessment for Diffusion-Based Novel View SynthesisInseong Choi, Siwoo Lee, Seung-Hun Nam, Soohwan SongCVPR 2026
