CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
Kaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai, Xintao Wang, Zinan Lin, Xuefei Ning, Jiwen Yu, Yu Wang, Xihui Liu
摘要
Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the task of cinematic video generation with decoupled scene context: given multiple images of a static environment, the goal is to synthesize high-quality videos featuring dynamic subject while preserving the underlying scene consistency and following a user-specified camera trajectory. We present CineScene, a framework that leverages implicit 3D-aware scene representation for cinematic video generation. Our key innovation is a novel context conditioning mechanism that injects 3D-aware features in an implicit way: By encoding scene images into visual representations through VGGT, CineScene injects spatial priors into a pretrained text-to-video generation model by additional context concatenation, enabling camera-controlled video synthesis with consistent scenes and dynamic subjects. To further enhance the model's robustness, we introduce a simple yet effective random-shuffling strategy for the input scene images during training. To address the lack of training data, we construct a scene-decoupled dataset with Unreal Engine 5, containing paired videos of scenes with and without dynamic subjects, panoramic images representing the underlying static scene, along with their camera trajectories. Experiments show that CineScene achieves state-of-the-art performance in scene-consistent cinematic video generation, handling large camera movements and demonstrating generalization across diverse environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
相关 Paper
- Recammaster: Camera-Controlled Generative Rendering From a Single VideoJianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang 等ICCV 2025 · 被引用 33 次
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry ContextJiaKui Hu, Jialun Liu, Liying Yang, Xinliang Zhang 等CVPR 2026 · 被引用 7 次
- 3D Scene Prompting for Scene-Consistent Camera-Controllable Video GenerationJoungBin Lee, Jaewoo Jung, Jisang Han, Takuya Narihira 等ICLR 2026 · 被引用 14 次
- 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsHeng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace 等NeurIPS 2024 · 被引用 76 次
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D ScenesYu Li, Menghan Xia, Gongye Liu, Jianhong Bai 等ICLR 2026 · 被引用 3 次
