Generalizable Implicit Motion Modeling for Video Frame Interpolation
Zujin Guo, Wei Li, Chen Change Loy
Abstract
Motion modeling is critical in flow-based Video Frame Interpolation (VFI). Existing paradigms either consider linear combinations of bidirectional flows or directly predict bilateral flows for given timestamps without exploring favorable motion priors, thus lacking the capability of effectively modeling spatiotemporal dynamics in real-world videos. To address this limitation, in this study, we introduce Generalizable Implicit Motion Modeling (GIMM), a novel and effective approach to motion modeling for VFI. Specifically, to enable GIMM as an effective motion modeling paradigm, we design a motion encoding pipeline to model spatiotemporal motion latent from bidirectional flows extracted from pre-trained flow estimators, effectively representing input-specific motion priors. Then, we implicitly predict arbitrary-timestep optical flows within two adjacent input frames via an adaptive coordinate-based neural network, with spatiotemporal coordinates and motion latent as inputs. Our GIMM can be easily integrated with existing flow-based VFI works by supplying accurately modeled motion. We show that GIMM performs better than the current state of the art on standard VFI benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69fd95bb-5297-4d00-ab99-3ecc98219591Cited by top-tier papers11
- Controllable Human-centric Keyframe Interpolation with Generative PriorZujin Guo, Size Wu, Zhongang Cai, Wei Li et al.NeurIPS 2025 · 6 citations
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion TransformersXinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng et al.CVPR 2026 · 4 citations
- Controllable Tracking-Based Video Frame InterpolationKarlis Martins Briedis, Abdelaziz Djelouah, Raphaël Ortiz, Markus Gross et al.SIGGRAPH 2025 · 4 citations
- SeeU: Seeing the Unseen World via 4D Dynamics-aware GenerationYu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang et al.CVPR 2026 · 3 citations
- Repurposing Pre-trained Video Diffusion Models for Event-based Video InterpolationJingxi Chen, Brandon Y. Feng, Haoming Cai, Tianfu Wang et al.CVPR 2025
Builds on28
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
- Channel Attention Is All You Need for Video Frame InterpolationMyungsub Choi, Heewon Kim, Bohyung Han, Ning Xu et al.AAAI 2020 · 362 citations
Related papers
- BiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform MotionsWonyong Seo, Jihyong Oh, Munchurl KimCVPR 2025
- VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame InterpolationChenyang Wu, Jiayi Fu, Chun-Le Guo, Shuhao Han et al.AAAI 2026 · 2 citations
- Sparse Global Matching for Video Frame Interpolation with Large MotionChunxu Liu, Guozhen Zhang, Rui Zhao, Limin WangCVPR 2024 · 17 citations
- Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Yijie Yu, Ling Yang, Chujun Qin et al.ACM MM 2024 · 10 citations
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
