Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single Image
Liao Shen, Xingyi Li, Huiqiang Sun, Juewen Peng, Ke Xian, Zhiguo Cao, Guosheng Lin
摘要
We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent perpetual views or struggle with long camera trajectories. To address these issues, it is essential to estimate the underlying 4D (including 3D geometry and scene motion) and fill in the occluded regions. To this end, we present Make-It-4D, a novel method that can generate a consistent long-term dynamic video from a single image. On the one hand, we utilize layered depth images (LDIs) to represent a scene, and they are then unprojected to form a feature point cloud. To animate the visual content, the feature point cloud is displaced based on the scene flow derived from motion estimation and the corresponding camera pose. Such 4D representation enables our method to maintain the global consistency of the generated dynamic video. On the other hand, we fill in the occluded regions by using a pre-trained diffusion model to inpaint and outpaint the input image. This enables our method to work under large camera motions. Benefiting from our design, our method can be training-free which saves a significant amount of training time. Experimental results demonstrate the effectiveness of our approach, which showcases compelling rendering results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Hierarchical Patch Diffusion Models for High-Resolution Video GenerationIvan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Sergey TulyakovCVPR 2024 · 被引用 14 次
- ElastoGen: 4D Generative ElastodynamicsYutao Feng, Yintong Shang, Xiang Feng, Lei Lan 等AAAI 2026 · 被引用 11 次
- SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body ManipulationMu Huang, Hui Wang, Kerui Ren, Linning Xu 等ICML 2026 · 被引用 3 次
- 3D-SceneDreamer: Text-Driven 3D-Consistent Scene GenerationSongchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma 等CVPR 2024 · 被引用 3 次
- SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One ModelYihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
相关 Paper
- Free4D: Tuning-Free 4D Scene Generation with Spatial-Temporal ConsistencyTianqi Liu, Zihao Huang, Zhaoxi Chen, Guangcong Wang 等ICCV 2025 · 被引用 3 次
- Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImagesIn-Hwan Jin, Haesoo Choo, Seong-Hun Jeong, Park Heemoon 等ICLR 2025
- Voyaging into Perpetual Dynamic Scenes from a Single ViewFengrui Tian, Tianjiao Ding, Jinqi Luo, Hancheng Min 等ICCV 2025 · 被引用 1 次
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu 等ICLR 2026 · 被引用 47 次
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImageFei Yin, Mallikarjun B. R., Chun-Han Yao, Rafal K. Mantiuk 等ICCV 2025 · 被引用 2 次
