SceneScape: Text-Driven Consistent Scene Generation
Rafail Fridman, Amit Abecasis, Yoni Kasten, Tali Dekel
摘要
We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such videos in an online fashion by combining the generative power of a pre-trained text-to-image model with the geometric priors learned by a pre-trained monocular depth prediction model. To tackle the pivotal challenge of achieving 3D consistency, i.e., synthesizing videos that depict geometrically-plausible scenes, we deploy an online test-time training to encourage the predicted depth map of the current frame to be geometrically consistent with the synthesized scene. The depth maps are used to construct a unified mesh representation of the scene, which is progressively constructed along the video generation process. In contrast to previous works, which are applicable only to limited domains, our method generates diverse scenes, such as walkthroughs in spaceships, caves, or ice castles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 被引用 439 次
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsLukas Höllein, Ang Cao, Andrew Owens, Justin Johnson 等ICCV 2023 · 被引用 292 次
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 等NeurIPS 2023 · 被引用 249 次
- Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct SupervisionAyush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov 等NeurIPS 2023 · 被引用 131 次
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera ControlZhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu 等NeurIPS 2024 · 被引用 131 次
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- Infinite Nature: Perpetual View Generation of Natural Scenes from a Single ImageAndrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely 等ICCV 2021 · 被引用 260 次
- Endless World: Real-Time 3D-Aware Long Video GenerationKe Zhang, Jiacong Xu, Yiqun Mei, Vishal M. PatelCVPR 2026 · 被引用 4 次
- Consistent depth of moving objects in videoZhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman 等SIGGRAPH 2021 · 被引用 26 次
- MoCa: Modeling Object Consistency for 3D Camera Control in Video GenerationZhijing Cheng, Xuancheng Zhang, Donglin Di, Chen Wei 等ICLR 2026
- 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsHeng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace 等NeurIPS 2024 · 被引用 76 次
