4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models
Heng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace, Aliaksandr Siarohin, Junli Cao, László A. Jeni, Sergey Tulyakov, Hsin-Ying Lee
Abstract
Existing dynamic scene generation methods mostly rely on distilling knowledge from pre-trained 3D generative models, which are typically fine-tuned on synthetic object datasets. As a result, the generated scenes are often object-centric and lack photorealism. To address these limitations, we introduce a novel pipeline designed for photorealistic text-to-4D scene generation, discarding the dependency on multi-view generative models and instead fully utilizing video generative models trained on diverse real-world datasets. Our method begins by generating a reference video using the video generation model. We then learn the canonical 3D representation of the video using a freeze-time video, delicately generated from the reference video. To handle inconsistencies in the freeze-time video, we jointly learn a per-frame deformation to model these imperfections. We then learn the temporal deformation based on the canonical representation to capture dynamic interactions in the reference video. The pipeline facilitates the generation of dynamic scenes with enhanced photorealism and structural integrity, viewable from multiple perspectives, thereby setting a new standard in 4D scene generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42103c69-9a27-4663-91dc-893d508967bfCited by top-tier papers3
- WorldScore: A Unified Evaluation Benchmark for World GenerationHaoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei et al.ICCV 2025 · 14 citations
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen et al.ICCV 2025 · 11 citations
- InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelsYifan Lu, Xuanchi Ren, Jiawei Yang, Tianchang Shen et al.ICCV 2025 · 10 citations
Builds on39
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D ScenesYu Li, Menghan Xia, Gongye Liu, Jianhong Bai et al.ICLR 2026 · 3 citations
- ShapeGen4D: Towards High Quality 4D Shape Generation from VideosJiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou et al.ICLR 2026 · 21 citations
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single VideoJixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan YangNeurIPS 2025 · 2 citations
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View ConsistencyYiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang et al.ICLR 2025
- Free4D: Tuning-Free 4D Scene Generation with Spatial-Temporal ConsistencyTianqi Liu, Zihao Huang, Zhaoxi Chen, Guangcong Wang et al.ICCV 2025 · 3 citations
