Playable Environments: Video Manipulation in Space and Time
Willi Menapace, Stéphane Lathuilière, Aliaksandr Siarohin, Christian Theobalt, Sergey Tulyakov, Vladislav Golyanik, Elisa Ricci
摘要
We present Playable Environments-a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actions are learnt in an unsupervised manner. The camera can be controlled to get the desired viewpoint. Our method builds an environment state for each frame, which can be manipulated by our proposed action mod-ule and decoded back to the image space with volumetric rendering. To support diverse appearances of objects, we extend neural radiance fields with style-based modulation. Our method trains on a collection of various monocular videos requiring only the estimated camera parameters and 2D object locations. To set a challenging benchmark, we in-troduce two large scale video datasets with significant cam-era movements. As evidenced by our experiments, playable environments enable several creative applications not at-tainable by prior video synthesis works, including playable 3D video generation, stylization and manipulation <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> willi-menapace.github.io/playable-environments-website.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh 等ICLR 2026 · 被引用 12 次
- Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single VideoHongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong WangCVPR 2024 · 被引用 9 次
- GameFactorly: Creating New Games with Generative Interactive VideosJiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper14
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoEdgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer 等ICCV 2021 · 被引用 617 次
- A Good Image Generator Is What You Need for High-Resolution Video SynthesisYu Tian, Jian Ren, Menglei Chai, Kyle Olszewski 等ICLR 2021 · 被引用 208 次
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn 等ICLR 2020 · 被引用 142 次
- GANcraft: Unsupervised 3D Neural Rendering of Minecraft WorldsZekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu LiuICCV 2021 · 被引用 131 次
相关 Paper
- Playable Video GenerationWilli Menapace, Stéphane Lathuilière, Sergey Tulyakov, Aliaksandr Siarohin 等CVPR 2021
- Editable free-viewpoint video using a layered neural representationJiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao 等SIGGRAPH 2021 · 被引用 80 次
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 被引用 316 次
- MonoNeRF: Learning Generalizable NeRFs from Monocular Videos without Camera PosesYang Fu, Ishan Misra, Xiaolong WangICML 2023 · 被引用 13 次
- STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural RenderingWentao Yuan, Zhaoyang Lv, Tanner Schmidt, Steven LovegroveCVPR 2021
