Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
Minghao Yin, Yukang Cao, Songyou Peng, Kai Han
Abstract
Generating high-quality 4D content from monocular videos-for applications such as digital humans and AR/VR-poses challenges in ensuring temporal and spatial consistency, preserving intricate details, and incorporating user guidance effectively. To overcome these challenges, we introduce Splat4D, a novel framework enabling high-fidelity 4D content generation from a monocular video. Splat4D achieves superior performance while maintaining faithful spatial-temporal coherence, by leveraging multi-view rendering, inconsistency identification, a video diffusion model, and an asymmetric U-Net for refinement. Through extensive evaluations on public benchmarks, Splat4D consistently demonstrates state-of-the-art performance across various metrics, underscoring the efficacy of our approach. Additionally, the versatility of Splat4D is validated in various applications such as text/image conditioned 4D generation, 4D human generation, and text-guided content editing, producing coherent outcomes following user instructions. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cf2dea6-a722-430a-8fd1-12862d110498Cited by top-tier papers3
- Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion TransformersMinghao Yin, Wenbo Hu, Jiale Xu, Ying Shan et al.CVPR 2026 · 3 citations
- VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied AgentsGeorge Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen et al.CVPR 2026
- Geometric-Photometric Event-based 3D Gaussian Ray TracingKai Kohyama, Yoshimitsu Aoki, Guillermo Gallego, Shintaro ShibaCVPR 2026
Builds on37
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- MVTokenFlow: High-quality 4D Content Generation using Multiview Token FlowHanzhuo Huang, Yuan Liu, Ge Zheng, Jiepeng Wang et al.ICLR 2025
- Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsHanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang et al.NeurIPS 2024 · 116 citations
- Mono4DGS-HDR: High Dynamic Range 4D Gaussian Splatting from Alternating-exposure Monocular VideosJinfeng Liu, Lingtong Kong, Mi Zhou, Jinwei Chen et al.ICLR 2026 · 3 citations
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsMichal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang et al.NeurIPS 2025 · 5 citations
- Sparse4DGS: Flow-Geometry Assisted 4D Gaussian Splatting for Dynamic Sparse View SynthesisDongdong Hu, Yang Zhou, Xiaofeng Huang, Haibing Yin et al.ACM MM 2025 · 3 citations
