Video Prediction via Example Guidance
Jingwei Xu, Huazhe Xu, Bingbing Ni, Xiaokang Yang, Trevor Darrell
Abstract
In video prediction tasks, one major challenge is to capture the multi-modal nature of future contents and dynamics. In this work, we propose a simple yet effective framework that can efficiently predict plausible future states. The key insight is that the potential distribution of a sequence could be approximated with analogous ones in a repertoire of training pool, namely, expert examples. By further incorporating a novel optimization scheme into the training procedure, plausible predictions can be sampled efficiently from distribution constructed from the retrieved examples. Meanwhile, our method could be seamlessly integrated with existing stochastic predictive models; significant enhancement is observed with comprehensive experiments in both quantitative and qualitative aspects. We also demonstrate the generalization ability to predict the motion of unseen class, i.e., without access to corresponding data during training phase. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 698a5d5a-5bb8-466f-9797-da6279fd4508Cited by top-tier papers5
- Where are you heading? Dynamic Trajectory Prediction with Expert Goal ExamplesHe Zhao, Richard P. WildesICCV 2021 · 74 citations
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.CVPR 2022 · 57 citations
- Multi-Objective Diverse Human Motion Prediction with Knowledge DistillationHengbo Ma, Jiachen Li, Ramtin Hosseini, Masayoshi Tomizuka et al.CVPR 2022 · 41 citations
- PastNet: Introducing Physical Inductive Biases for Spatio-temporal Video PredictionHao Wu, Fan Xu, Chong Chen, Xian-Sheng Hua et al.ACM MM 2024 · 36 citations
- A Frame is Worth One Token: Efficient Generative World Modeling with Delta TokensTommie Kerssies, Gabriele Berton, Ju He, Qihang Yu et al.CVPR 2026 · 8 citations
Builds on4
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
- Deep Kinematics Analysis for Monocular 3D Human Pose EstimationJingwei Xu, Zhenbo Yu, Bingbing Ni, Jiancheng Yang et al.CVPR 2020
Related papers
- PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and PlanningAngel Villar-Corrales, Sven BehnkeICML 2025
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould et al.ICCV 2021 · 44 citations
- Diverse Video Generation using a Gaussian Process TriggerGaurav Shrivastava, Abhinav ShrivastavaICLR 2021 · 22 citations
- Optimizing Video Prediction via Video Frame InterpolationYue Wu, Qiang Wen, Qifeng ChenCVPR 2022 · 47 citations
- Forecasting Characteristic 3D Poses of Human ActionsChristian Diller, Thomas A. Funkhouser, Angela DaiCVPR 2022 · 22 citations
