Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation
Sunghyun Park, Kangyeol Kim, Junsoo Lee, Jaegul Choo, Joonseok Lee, Sookyung Kim, Edward Choi
摘要
Video generation models often operate under the assumption of fixed frame rates, which leads to suboptimal performance when it comes to handling flexible frame rates (e.g., increasing the frame rate of the more dynamic portion of the video as well as handling missing video frames). To resolve the restricted nature of existing video generation models' ability to handle arbitrary timesteps, we propose continuous-time video generation by combining neural ODE (Vid-ODE) with pixel-level video processing techniques. Using ODE-ConvGRU as an encoder, a convolutional version of the recently proposed neural ODE, which enables us to learn continuous-time dynamics, Vid-ODE can learn the spatio-temporal dynamics of input videos of flexible frame rates. The decoder integrates the learned dynamics function to synthesize video frames at any given timesteps, where the pixel-level composition technique is used to maintain the sharpness of individual frames. With extensive experiments on four real-world video datasets, we verify that the proposed Vid-ODE outperforms state-of-the-art approaches under various video generation settings, both within the trained time range (interpolation) and beyond the range (extrapolation). To the best of our knowledge, Vid-ODE is the first work successfully performing continuous-time video generation using real-world videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 被引用 167 次
- Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One ModelHao Wu, Yuxuan Liang, Wei Xiong, Zhengyang Zhou 等AAAI 2024 · 被引用 58 次
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman 等NeurIPS 2023 · 被引用 36 次
- Motion-Aware Dynamic Architecture for Efficient Frame InterpolationMyungsub Choi, Suyoung Lee, Heewon Kim, Kyoung Mu LeeICCV 2021 · 被引用 26 次
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 被引用 20 次
它引用的顶会 Paper3
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier 等ICML 2020 · 被引用 166 次
- Unsupervised Video Interpolation Using Cycle ConsistencyFitsum A. Reda, Deqing Sun, Aysegul Dundar, Mohammad Shoeybi 等ICCV 2019 · 被引用 93 次
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang 等ICCV 2019 · 被引用 90 次
相关 Paper
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal 等ICCV 2023 · 被引用 3 次
- Generative Video Bi-FlowChen Liu, Tobias RitschelICCV 2025 · 被引用 2 次
- Neural Jump Ordinary Differential Equations: Consistent Continuous-Time Prediction and FilteringCalypso Herrera, Florian Krach, Josef TeichmannICLR 2021 · 被引用 43 次
- MoStGAN-V: Video Generation with Temporal Motion StylesXiaoqian Shen, Xiang Li, Mohamed ElhoseinyCVPR 2023
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao 等NeurIPS 2024 · 被引用 62 次
