Diverse Video Generation using a Gaussian Process Trigger
Gaurav Shrivastava, Abhinav Shrivastava
Abstract
Generating future frames given a few context (or past) frames is a challenging task. It requires modeling the temporal coherence of videos and multi-modality in terms of diversity in the potential future states. Current variational approaches for video generation tend to marginalize over multi-modal future outcomes. Instead, we propose to explicitly model the multi-modality in the future outcomes and leverage it to sample diverse futures. Our approach, Diverse Video Generator, uses a Gaussian Process (GP) to learn priors on future states given the past and maintains a probability distribution over possible futures given a particular sample. In addition, we leverage the changes in this distribution overtime to control the sampling of diverse future states by estimating the end of on-going sequences. That is, we use the variance of GP over the output function space to trigger a change in an action sequence. We achieve state-of-the-art results on diverse future frame generation in terms of reconstruction quality and diversity of the generated sequences. Webpage -
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1bde9b1-feb3-4897-863f-e31fa3d6d434Cited by top-tier papers3
- Video Dynamics Prior: An Internal Learning Approach for Robust Video EnhancementsGaurav Shrivastava, Ser Nam Lim, Abhinav ShrivastavaNeurIPS 2023 · 14 citations
- Video Prediction by Modeling Videos as Continuous Multi-Dimensional ProcessesGaurav Shrivastava, Abhinav ShrivastavaCVPR 2024 · 4 citations
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
Builds on3
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
Related papers
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould et al.ICCV 2021 · 44 citations
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
- Video Instance Segmentation Tracking With a Modified VAE ArchitectureChung-Ching Lin, Ying Hung, Rogério Feris, Linglin HeCVPR 2020
- Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single SampleShir Gur, Sagie Benaim, Lior WolfNeurIPS 2020 · 84 citations
