Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D Synthesis
Bowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao, Dong Chen, Baining Guo
Abstract
In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appearance, and motion. We address these challenges by introducing a Direct 4DMesh-to-GS Variation Field VAE that directly encodes canonical Gaussian Splats (GS) and their temporal variations from 3D animation data without per-instance fitting, and compresses high-dimensional animations into a compact latent space. Building upon this efficient representation, we train a Gaussian Variation Field diffusion model with temporal-aware Diffusion Transformer conditioned on input videos and canonical GS. Trained on carefully-curated animatable 3D objects from the Objaverse dataset, our model demonstrates superior generation quality compared to existing methods. It also exhibits remarkable generalization to in-the-wild video inputs despite being trained exclusively on synthetic data, paving the way for generating high-quality animated 3D content. Project page: https://gvfdiffusion.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31797107-01df-41ec-b031-97b408b38f9fCited by top-tier papers13
- ShapeGen4D: Towards High Quality 4D Shape Generation from VideosJiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou et al.ICLR 2026 · 21 citations
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
- Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular VideoZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.CVPR 2026 · 13 citations
- PerpetualWonder: Long-horizon Action-conditioned 4D Scene GenerationJiahao Zhan, Zizhang Li, Hong-Xing Yu, Jiajun WuCVPR 2026 · 9 citations
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character GenerationMiaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li et al.CVPR 2026 · 6 citations
Builds on60
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- L4GM: Large 4D Gaussian Reconstruction ModelJiawei Ren, Cheng Xie, Ashkan Mirzaei, Hanxue Liang et al.NeurIPS 2024 · 173 citations
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene GenerationPanwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao et al.CVPR 2026
- Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion ModelsHuan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler et al.CVPR 2024
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View ConsistencyYiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang et al.ICLR 2025
- ActionMesh: Animated 3D Mesh Generation with Temporal 3D DiffusionRemy Sabathier, David Novotný, Niloy J. Mitra, Tom MonnierCVPR 2026 · 16 citations
