Efficient Video Prediction via Sparsely Conditioned Flow Matching
Aram Davtyan, Sepehr Sameni, Paolo Favaro
Abstract
We introduce a novel generative model for video prediction based on latent flow matching, an efficient alternative to diffusion-based models. In contrast to prior work, we keep the high costs of modeling the past during training and inference at bay by conditioning only on a small random set of past frames at each integration step of the image generation process. Moreover, to enable the generation of high-resolution videos and to speed up the training, we work in the latent space of a pretrained VQGAN. Finally, we propose to approximate the initial condition of the flow ODE with the previous noisy frame. This allows to reduce the number of integration steps and hence, speed up the sampling at inference time. We call our model Random frame conditioned flow Integration for VidEo pRediction, or, in short, RIVER. We show that RIVER achieves superior or on par performance compared to prior work on common video prediction benchmarks, while requiring an order of magnitude fewer computational resources. Project website: https://araachie.github.io/river.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a840550a-bff1-447d-a059-644d7a47ccc5Cited by top-tier papers24
- Probabilistic Forecasting with Stochastic Interpolants and Föllmer ProcessesYifan Chen, Mark Goldstein, Mengjian Hua, Michael S. Albergo et al.ICML 2024 · 53 citations
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesJin Wang, Yao Lai, Aoxue Li, Shifeng Zhang et al.NeurIPS 2025 · 45 citations
- Dynamic Conditional Optimal Transport through Simulation-Free FlowsGavin Kerrigan, Giosue Migliorini, Padhraic SmythNeurIPS 2024 · 36 citations
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh et al.ICLR 2026 · 12 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Generative Video Bi-FlowChen Liu, Tobias RitschelICCV 2025 · 2 citations
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo et al.ICLR 2025
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li et al.ICLR 2024 · 34 citations
