Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation
Sunghyun Park, Kangyeol Kim, Junsoo Lee, Jaegul Choo, Joonseok Lee, Sookyung Kim, Edward Choi
Abstract
Video generation models often operate under the assumption of fixed frame rates, which leads to suboptimal performance when it comes to handling flexible frame rates (e.g., increasing the frame rate of the more dynamic portion of the video as well as handling missing video frames). To resolve the restricted nature of existing video generation models' ability to handle arbitrary timesteps, we propose continuous-time video generation by combining neural ODE (Vid-ODE) with pixel-level video processing techniques. Using ODE-ConvGRU as an encoder, a convolutional version of the recently proposed neural ODE, which enables us to learn continuous-time dynamics, Vid-ODE can learn the spatio-temporal dynamics of input videos of flexible frame rates. The decoder integrates the learned dynamics function to synthesize video frames at any given timesteps, where the pixel-level composition technique is used to maintain the sharpness of individual frames. With extensive experiments on four real-world video datasets, we verify that the proposed Vid-ODE outperforms state-of-the-art approaches under various video generation settings, both within the trained time range (interpolation) and beyond the range (extrapolation). To the best of our knowledge, Vid-ODE is the first work successfully performing continuous-time video generation using real-world videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb9ce3c1-b803-42ad-b25d-abbca5351038Cited by top-tier papers22
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
- Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One ModelHao Wu, Yuxuan Liang, Wei Xiong, Zhengyang Zhou et al.AAAI 2024 · 58 citations
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman et al.NeurIPS 2023 · 36 citations
- Motion-Aware Dynamic Architecture for Efficient Frame InterpolationMyungsub Choi, Suyoung Lee, Heewon Kim, Kyoung Mu LeeICCV 2021 · 26 citations
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 20 citations
Builds on3
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
- Unsupervised Video Interpolation Using Cycle ConsistencyFitsum A. Reda, Deqing Sun, Aysegul Dundar, Mohammad Shoeybi et al.ICCV 2019 · 93 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
Related papers
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal et al.ICCV 2023 · 3 citations
- Generative Video Bi-FlowChen Liu, Tobias RitschelICCV 2025 · 2 citations
- Neural Jump Ordinary Differential Equations: Consistent Continuous-Time Prediction and FilteringCalypso Herrera, Florian Krach, Josef TeichmannICLR 2021 · 43 citations
- MoStGAN-V: Video Generation with Temporal Motion StylesXiaoqian Shen, Xiang Li, Mohamed ElhoseinyCVPR 2023
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao et al.NeurIPS 2024 · 62 citations
