Playable Environments: Video Manipulation in Space and Time
Willi Menapace, Stéphane Lathuilière, Aliaksandr Siarohin, Christian Theobalt, Sergey Tulyakov, Vladislav Golyanik, Elisa Ricci
Abstract
We present Playable Environments-a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actions are learnt in an unsupervised manner. The camera can be controlled to get the desired viewpoint. Our method builds an environment state for each frame, which can be manipulated by our proposed action mod-ule and decoded back to the image space with volumetric rendering. To support diverse appearances of objects, we extend neural radiance fields with style-based modulation. Our method trains on a collection of various monocular videos requiring only the estimated camera parameters and 2D object locations. To set a challenging benchmark, we in-troduce two large scale video datasets with significant cam-era movements. As evidenced by our experiments, playable environments enable several creative applications not at-tainable by prior video synthesis works, including playable 3D video generation, stylization and manipulation <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> willi-menapace.github.io/playable-environments-website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddd2f30b-2d73-4c21-8de6-70a3ec6145b3Cited by top-tier papers11
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh et al.ICLR 2026 · 12 citations
- Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single VideoHongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong WangCVPR 2024 · 9 citations
- GameFactorly: Creating New Games with Generative Interactive VideosJiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan et al.ICCV 2025 · 7 citations
Builds on14
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoEdgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer et al.ICCV 2021 · 617 citations
- A Good Image Generator Is What You Need for High-Resolution Video SynthesisYu Tian, Jian Ren, Menglei Chai, Kyle Olszewski et al.ICLR 2021 · 208 citations
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
- GANcraft: Unsupervised 3D Neural Rendering of Minecraft WorldsZekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu LiuICCV 2021 · 131 citations
Related papers
- Playable Video GenerationWilli Menapace, Stéphane Lathuilière, Sergey Tulyakov, Aliaksandr Siarohin et al.CVPR 2021
- Editable free-viewpoint video using a layered neural representationJiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao et al.SIGGRAPH 2021 · 80 citations
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
- MonoNeRF: Learning Generalizable NeRFs from Monocular Videos without Camera PosesYang Fu, Ishan Misra, Xiaolong WangICML 2023 · 13 citations
- STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural RenderingWentao Yuan, Zhaoyang Lv, Tanner Schmidt, Steven LovegroveCVPR 2021
