MOSO: Decomposing MOtion, Scene and Object for Video Prediction
Mingzhen Sun, Weining Wang, Xinxin Zhu, Jing Liu
摘要
Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a twostage MOtion, Scene and Object decomposition framework (MOSO) 1 for video prediction, consisting of MOSO-VQVAE and MOSO-Transformer. In the first stage, MOSO-VQVAE decomposes a previous video clip into the motion, scene and object components, and represents them as distinct groups of discrete tokens. Then, in the second stage, MOSO-Transformer predicts the object and scene tokens of the subsequent video clip based on the previous tokens and adds dynamic motion at the token level to the generated object and scene tokens. Our framework can be easily extended to unconditional video generation and video frame interpolation tasks. Experimental results demonstrate that our method achieves new state-of-the-art performance on five challenging benchmarks for video prediction and unconditional video generation: BAIR, RoboNet, KTH, KITTI and UCF101. In addition, MOSO can produce realistic videos by combining objects and scenes from different videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 被引用 20 次
- Generative Pre-trained Autoregressive Diffusion TransformerYuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu 等NeurIPS 2025 · 被引用 19 次
- GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodERMingzhen Sun, Weining Wang, Zihan Qin, Jiahui Sun 等NeurIPS 2023 · 被引用 7 次
- Motion Graph Unleashed: A Novel Approach to Video PredictionYiqi Zhong, Luming Liang, Bohan Tang, Ilya Zharkov 等NeurIPS 2024 · 被引用 7 次
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun 等ACM MM 2024 · 被引用 4 次
它引用的顶会 Paper16
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 被引用 434 次
- PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersXiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen 等AAAI 2023 · 被引用 281 次
相关 Paper
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin 等CVPR 2025
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu 等ICML 2024 · 被引用 94 次
- WALDO: Future Video Synthesis using Object Layer Decomposition and Parametric Flow PredictionGuillaume Le Moing, Jean Ponce, Cordelia SchmidICCV 2023 · 被引用 7 次
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 等ICML 2026 · 被引用 9 次
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 被引用 37 次
