SceneScape: Text-Driven Consistent Scene Generation
Rafail Fridman, Amit Abecasis, Yoni Kasten, Tali Dekel
Abstract
We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such videos in an online fashion by combining the generative power of a pre-trained text-to-image model with the geometric priors learned by a pre-trained monocular depth prediction model. To tackle the pivotal challenge of achieving 3D consistency, i.e., synthesizing videos that depict geometrically-plausible scenes, we deploy an online test-time training to encourage the predicted depth map of the current frame to be geometrically consistent with the synthesized scene. The depth maps are used to construct a unified mesh representation of the scene, which is progressively constructed along the video generation process. In contrast to previous works, which are applicable only to limited domains, our method generates diverse scenes, such as walkthroughs in spaceships, caves, or ice castles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers64
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsLukas Höllein, Ang Cao, Andrew Owens, Justin Johnson et al.ICCV 2023 · 292 citations
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang et al.NeurIPS 2023 · 249 citations
- Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct SupervisionAyush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov et al.NeurIPS 2023 · 131 citations
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera ControlZhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu et al.NeurIPS 2024 · 131 citations
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Infinite Nature: Perpetual View Generation of Natural Scenes from a Single ImageAndrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely et al.ICCV 2021 · 260 citations
- Endless World: Real-Time 3D-Aware Long Video GenerationKe Zhang, Jiacong Xu, Yiqun Mei, Vishal M. PatelCVPR 2026 · 4 citations
- Consistent depth of moving objects in videoZhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman et al.SIGGRAPH 2021 · 26 citations
- MoCa: Modeling Object Consistency for 3D Camera Control in Video GenerationZhijing Cheng, Xuancheng Zhang, Donglin Di, Chen Wei et al.ICLR 2026
- 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsHeng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace et al.NeurIPS 2024 · 76 citations
