Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels
Yikai Wang, Xinzhou Wang, Zilong Chen, Zhengyi Wang, Fuchun Sun, Jun Zhu
Abstract
Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D, a novel reconstruction model that excels in accurately reconstructing 4D (i.e., sequential 3D) representations from single generated videos, addressing challenges associated with non-rigidity and frame distortion. This capability is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. At the core of Vidu4D is our proposed Dynamic Gaussian Surfels (DGS) technique. DGS optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. This transformation enables a precise depiction of motion and deformation over time. To preserve the structural integrity of surface-aligned Gaussian surfels, we design the warped-state geometric regularization based on continuous warping fields for estimating normals. Additionally, we learn refinements on rotation and scaling parameters of Gaussian surfels, which greatly alleviates texture flickering during the warping process and enhances the capture of fine-grained appearance details. Vidu4D also contains a novel initialization state that provides a proper start for the warping fields in DGS. Equipping Vidu4D with an existing video generative model, the overall framework demonstrates high-fidelity text-to-4D generation in both appearance and geometry.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 015fdb50-ac40-40cc-892b-569580235f8bCited by top-tier papers22
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionYixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu et al.ICLR 2026 · 34 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- ActionMesh: Animated 3D Mesh Generation with Temporal 3D DiffusionRemy Sabathier, David Novotný, Niloy J. Mitra, Tom MonnierCVPR 2026 · 16 citations
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen et al.ICCV 2025 · 11 citations
Builds on66
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- 4DSurf: High-Fidelity Dynamic Scene Surface ReconstructionRenjie Wu, Hongdong Li, José M. Álvarez, Miaomiao LiuCVPR 2026
- 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D GenerationMengmeng Liu, Jiuming Liu, Yunpeng Zhang, Jiangtao Li et al.AAAI 2026 · 2 citations
- Fused View-Time Attention and Feedforward Reconstruction for 4D Scene GenerationChaoyang Wang, Ashkan Mirzaei, Vidit Goel, Willi Menapace et al.NeurIPS 2025 · 12 citations
- Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D SynthesisBowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang et al.ICCV 2025 · 4 citations
- Deblur4DGS: 4D Gaussian Splatting from Blurry Monocular VideoRenlong Wu, Zhilu Zhang, Mingyang Chen, Zifei Yan et al.AAAI 2026 · 17 citations
