GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
Xingyilang Yin, Qi Zhang, Jiahao Chang, Ying Feng, Qingnan Fan, Xi Yang, Chi-Man Pun, Huaqi Zhang, Xiaodong Cun
Abstract
3D reconstruction and novel view synthesis (NVS) are fundamental tasks in computer vision and graphics, with wide-ranging real-world applications in virtual reality, autonomous driving, and robotics. Recently, 3D Gaussian Splatting (3DGS) (Kerbl et al., 2023) has achieved impressive results in both reconstruction quality and rendering efficiency when dense input views are available. However, its performance degrades significantly in sparse-view settings, where limited viewpoint information leads to under-constrained 3D representations. In such cases, 3DGS often suffers from severe artifacts, including distorted geometric structures and incomplete reconstructions, particularly in less-observed regions or extreme novel viewpoints. These limitations hinder its applicability in real-world scenarios where acquiring dense multi-view data is challenging. To alleviate the limitations, some previous regularization methods have been proposed to introduce additional constraints into the 3DGS optimization process, such as monocular depth (Li et al., 2024; Zhu et al., 2024) , frequency smoothness (Zhang et al., 2024) , and random dropout (Xu et al., 2025) . While these approaches can help prevent 3DGS representations from overfitting to sparse input views, they often remain sensitive to noise and yield only marginal improvements in NVS rendering quality. Inspired by the success of ReconFusion (Wu et al., 2024) , which introduces diffusion model into NeRF (Mildenhall et al., 2020) optimization, more recent studies (Liu et al., 2024b;a; Wu et al., 2025a;b) explore incorporating 3DGS optimization with powerful generative priors from diffusion models, which are trained on internet-scale data. These strong priors enable the correction of spurious geometry or the inpainting of plausible content in novel views. However, a key challenge still remains: maintaining visual and 3D consistency between the generated and original input images, especially when the novel views are far from the observed inputs. Meanwhile, recent advances in controllable video generation have demonstrated the effectiveness of incorporating various conditional signals (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4db7c88a-a7f0-4eb7-b3ad-0be8f54b7585Cited by top-tier papers8
- G4Splat: Geometry-Guided Gaussian Splatting with Generative PriorJunfeng Ni, Yixin Chen, Zhifei Yang, Yu Liu et al.ICLR 2026 · 10 citations
- GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video GeneratorLiyuan Zhu, Manjunath Narayana, Michal Stary, Will Hutchcroft et al.CVPR 2026 · 7 citations
- Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video DiffusionTing-Hsuan Chen, Ying-Huan Chen, Tao Tu, Jie-Ying Lee et al.CVPR 2026 · 2 citations
- HAD: Hallucination-Aware Diffusion Priors for 3D ReconstructionXi Liu, Weiwei Sun, Zhou Ren, Chris Broaddus et al.CVPR 2026 · 1 citation
- CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single ImageYizheng Song, Yiyu Zhuang, Qipeng Xu, Haixiang Wang et al.CVPR 2026 · 1 citation
Builds on41
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsYuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang et al.NeurIPS 2024 · 126 citations
- 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion PriorsXi Liu, Chaoyi Zhou, Siyu HuangNeurIPS 2024 · 127 citations
- GeoQuery: Geometry-Query Diffusion for Sparse-View ReconstructionXiao Cao, Yuze Li, Youmin Zhang, Jiayu Song et al.SIGGRAPH 2026
- FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage TrainingRuihong Yin, Vladimir Yugay, Yue Li, Sezer Karaoglu et al.NeurIPS 2024 · 29 citations
- DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View ConsistencyTengfei Xiao, Yue Wu, Zhigang Gao, Yongzhe Yuan et al.AAAI 2026
