Compositional Video Prediction
Yufei Ye, Maneesh Singh, Abhinav Gupta, Shubham Tulsiani
摘要
We present an approach for pixel-level future prediction given an input image of a scene. We observe that a scene is comprised of distinct entities that undergo motion and present an approach that operationalizes this insight. We implicitly predict future states of independent entities while reasoning about their interactions, and compose future video frames using these predicted states. We overcome the inherent multi-modality of the task using a global trajectory-level latent random variable, and show that this allows us to sample diverse and plausible futures. We empirically validate our approach against alternate representations and ways of incorporating multi-modality. We examine two datasets, one comprising of stacked objects that may fall, and the other containing videos of humans performing activities in a gym, and show that our approach allows realistic stochastic video prediction across these diverse settings. See project website (https://judyye.github.io/CVP/) for video predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Infinite Nature: Perpetual View Generation of Natural Scenes from a Single ImageAndrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely 等ICCV 2021 · 被引用 260 次
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 被引用 201 次
- Causal Discovery in Physical Systems from VideosYunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox 等NeurIPS 2020 · 被引用 133 次
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski 等NeurIPS 2023 · 被引用 106 次
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo 等NeurIPS 2021 · 被引用 90 次
相关 Paper
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesGabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht 等ICLR 2026 · 被引用 10 次
- Video Prediction via Example GuidanceJingwei Xu, Huazhe Xu, Bingbing Ni, Xiaokang Yang 等ICML 2020 · 被引用 18 次
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn 等ICLR 2020 · 被引用 142 次
- Learning Disentangled Representations of Videos with Missing DataArmand Comas Massague, Chi Zhang, Zlatan Feric, Octavia I. Camps 等NeurIPS 2020 · 被引用 17 次
- MOSO: Decomposing MOtion, Scene and Object for Video PredictionMingzhen Sun, Weining Wang, Xinxin Zhu, Jing LiuCVPR 2023
