Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single Image
Liao Shen, Xingyi Li, Huiqiang Sun, Juewen Peng, Ke Xian, Zhiguo Cao, Guosheng Lin
Abstract
We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent perpetual views or struggle with long camera trajectories. To address these issues, it is essential to estimate the underlying 4D (including 3D geometry and scene motion) and fill in the occluded regions. To this end, we present Make-It-4D, a novel method that can generate a consistent long-term dynamic video from a single image. On the one hand, we utilize layered depth images (LDIs) to represent a scene, and they are then unprojected to form a feature point cloud. To animate the visual content, the feature point cloud is displaced based on the scene flow derived from motion estimation and the corresponding camera pose. Such 4D representation enables our method to maintain the global consistency of the generated dynamic video. On the other hand, we fill in the occluded regions by using a pre-trained diffusion model to inpaint and outpaint the input image. This enables our method to work under large camera motions. Benefiting from our design, our method can be training-free which saves a significant amount of training time. Experimental results demonstrate the effectiveness of our approach, which showcases compelling rendering results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c938d572-797c-4334-9aff-778fb053f32cCited by top-tier papers11
- Hierarchical Patch Diffusion Models for High-Resolution Video GenerationIvan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Sergey TulyakovCVPR 2024 · 14 citations
- ElastoGen: 4D Generative ElastodynamicsYutao Feng, Yintong Shang, Xiang Feng, Lei Lan et al.AAAI 2026 · 11 citations
- SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body ManipulationMu Huang, Hui Wang, Kerui Ren, Linning Xu et al.ICML 2026 · 3 citations
- 3D-SceneDreamer: Text-Driven 3D-Consistent Scene GenerationSongchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma et al.CVPR 2024 · 3 citations
- SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelYihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao et al.ACM MM 2024 · 2 citations
Builds on25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
Related papers
- Free4D: Tuning-Free 4D Scene Generation with Spatial-Temporal ConsistencyTianqi Liu, Zihao Huang, Zhaoxi Chen, Guangcong Wang et al.ICCV 2025 · 3 citations
- Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImagesIn-Hwan Jin, Haesoo Choo, Seong-Hun Jeong, Park Heemoon et al.ICLR 2025
- Voyaging into Perpetual Dynamic Scenes from a Single ViewFengrui Tian, Tianjiao Ding, Jinqi Luo, Hancheng Min et al.ICCV 2025 · 1 citation
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImageFei Yin, Mallikarjun B. R., Chun-Han Yao, Rafal K. Mantiuk et al.ICCV 2025 · 2 citations
