Disentangling Propagation and Generation for Video Prediction
Hang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang, Fisher Yu, Trevor Darrell
摘要
A dynamic scene has two types of elements: those that move fluidly and can be predicted from previous frames, and those which are disoccluded (exposed) and cannot be extrapolated. Prior approaches to video prediction typically learn either to warp or to hallucinate future pixels, but not both. In this paper, we describe a computational model for high-fidelity video prediction which disentangles motion-specific propagation from motion-agnostic generation. We introduce a confidence-aware warping operator which gates the output of pixel predictions from a flow predictor for non-occluded regions and from a context encoder for occluded regions. Moreover, in contrast to prior works where confidence is jointly learned with flow and appearance using a single network, we compute confidence after a warping step, and employ a separate network to inpaint exposed regions. Empirical results on both synthetic and real datasets show that our disentangling approach provides better occlusion maps and produces both sharper and more realistic predictions compared to strong baselines. * Equal contribution. † Work was done while Q.C. and R.W. were at Berkeley DeepDrive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier 等ICML 2020 · 被引用 166 次
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 被引用 117 次
- CCVS: Context-aware Controllable Video SynthesisGuillaume Le Moing, Jean Ponce, Cordelia SchmidNeurIPS 2021 · 被引用 98 次
- Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential EquationSunghyun Park, Kangyeol Kim, Junsoo Lee, Jaegul Choo 等AAAI 2021 · 被引用 62 次
相关 Paper
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu 等AAAI 2024 · 被引用 13 次
- Future Video Synthesis With Object Motion PredictionYue Wu, Rongrong Gao, Jaesik Park, Qifeng ChenCVPR 2020
- SLAMP: Stochastic Latent Appearance and Motion PredictionAdil Kaan Akan, Erkut Erdem, Aykut Erdem, Fatma GüneyICCV 2021 · 被引用 46 次
- VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and EditingJuan Luis Gonzalez, Xu Yao, Alex Whelan, Kyle Olszewski 等CVPR 2025
