Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
Minghao Yin, Yukang Cao, Songyou Peng, Kai Han
摘要
Generating high-quality 4D content from monocular videos-for applications such as digital humans and AR/VR-poses challenges in ensuring temporal and spatial consistency, preserving intricate details, and incorporating user guidance effectively. To overcome these challenges, we introduce Splat4D, a novel framework enabling high-fidelity 4D content generation from a monocular video. Splat4D achieves superior performance while maintaining faithful spatial-temporal coherence, by leveraging multi-view rendering, inconsistency identification, a video diffusion model, and an asymmetric U-Net for refinement. Through extensive evaluations on public benchmarks, Splat4D consistently demonstrates state-of-the-art performance across various metrics, underscoring the efficacy of our approach. Additionally, the versatility of Splat4D is validated in various applications such as text/image conditioned 4D generation, 4D human generation, and text-guided content editing, producing coherent outcomes following user instructions. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion TransformersMinghao Yin, Wenbo Hu, Jiale Xu, Ying Shan 等CVPR 2026 · 被引用 3 次
- VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied AgentsGeorge Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen 等CVPR 2026
- Geometric-Photometric Event-based 3D Gaussian Ray TracingKai Kohyama, Yoshimitsu Aoki, Guillermo Gallego, Shintaro ShibaCVPR 2026
它引用的顶会 Paper37
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
相关 Paper
- MVTokenFlow: High-quality 4D Content Generation using Multiview Token FlowHanzhuo Huang, Yuan Liu, Ge Zheng, Jiepeng Wang 等ICLR 2025
- Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsHanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang 等NeurIPS 2024 · 被引用 116 次
- Mono4DGS-HDR: High Dynamic Range 4D Gaussian Splatting from Alternating-exposure Monocular VideosJinfeng Liu, Lingtong Kong, Mi Zhou, Jinwei Chen 等ICLR 2026 · 被引用 3 次
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsMichal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang 等NeurIPS 2025 · 被引用 5 次
- Sparse4DGS: Flow-Geometry Assisted 4D Gaussian Splatting for Dynamic Sparse View SynthesisDongdong Hu, Yang Zhou, Xiaofeng Huang, Haibing Yin 等ACM MM 2025 · 被引用 3 次
