WALDO: Future Video Synthesis using Object Layer Decomposition and Parametric Flow Prediction
Guillaume Le Moing, Jean Ponce, Cordelia Schmid
Abstract
This paper presents WALDO (WArping Layer-Decomposed Objects), a novel approach to the prediction of future video frames from past ones. Individual images are decomposed into multiple layers combining object masks and a small set of control points. The layer structure is shared across all frames in each video to build dense inter-frame connections. Complex scene motions are modeled by combining parametric geometric transformations associated with individual layers, and video synthesis is broken down into discovering the layers associated with past frames, predicting the corresponding transformations for upcoming ones and warping the associated object regions accordingly, and filling in the remaining image parts. Extensive experiments on multiple benchmarks including urban videos (Cityscapes and KITTI) and videos featuring nonrigid motions (UCF-Sports and H3.6M), show that our method consistently outperforms the state of the art by a significant margin in every case. Code, pretrained models, and video samples synthesized by our approach can be found in the project webpage. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2f94223-6661-4fe2-ab10-3a491834102aCited by top-tier papers1
Ask how each one uses itBuilds on38
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
Related papers
- Future Video Synthesis With Object Motion PredictionYue Wu, Rongrong Gao, Jaesik Park, Qifeng ChenCVPR 2020
- MOSO: Decomposing MOtion, Scene and Object for Video PredictionMingzhen Sun, Weining Wang, Xinxin Zhu, Jing LiuCVPR 2023
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
- Optimizing Video Prediction via Video Frame InterpolationYue Wu, Qiang Wen, Qifeng ChenCVPR 2022 · 47 citations
- Objects do not disappear: Video object detection by single-frame object location anticipationXin Liu, Fatemeh Karimi Nejadasl, Jan C. van Gemert, Olaf Booij et al.ICCV 2023 · 10 citations
