ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
Duolikun Danier, Fan Zhang, David Bull
Abstract
Video frame interpolation (VFI) is currently a very active research topic, with applications spanning computer vision, post production and video encoding. VFI can be extremely challenging, particularly in sequences containing large motions, occlusions or dynamic textures, where existing approaches fail to offer perceptually robust inter-polation performance. In this context, we present a novel deep learning based VFI method, ST-MFNet, based on a Spatio-Temporal Multi-Flow architecture. ST-MFNet employs a new multi-scale multi-flow predictor to estimate many-to-one intermediate flows, which are combined with conventional one-to-one optical flows to capture both large and complex motions. In order to enhance interpolation performance for various textures, a 3D CNN is also employed to model the content dynamics over an extended temporal window. Moreover, ST-MFNet has been trained within an ST-GAN framework, which was originally developedfor texture synthesis, with the aim of further improving perceptual interpolation quality. Our approach has been comprehensively evaluated - compared with fourteen state-of-the-art VFI algorithms - clearly demonstrating that ST-MFNet consistently outperforms these benchmarks on var-ied and representative test datasets, with significant gains up to 1.09dB in PSNR for cases including large motions and dynamic textures. Our source code is available at https://github.com/danielism97/ST-MFNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68e1fe14-dc50-40b8-bf7d-386bcac231ceCited by top-tier papers18
- LDMVFI: Video Frame Interpolation with Latent Diffusion ModelsDuolikun Danier, Fan Zhang, David BullAAAI 2024 · 115 citations
- Perception-Oriented Video Frame Interpolation via Asymmetric BlendingGuangyang Wu, Xin Tao, Changlin Li, Wenyi Wang et al.CVPR 2024 · 16 citations
- Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Yijie Yu, Ling Yang, Chujun Qin et al.ACM MM 2024 · 10 citations
- Kernel-Based Frame Interpolation for Spatio-Temporally Adaptive RenderingKarlis Martins Briedis, Abdelaziz Djelouah, Raphaël Ortiz, Mark Meyer et al.SIGGRAPH 2023 · 9 citations
- Video Object Segmentation-aware Video Frame InterpolationJun-Sang Yoo, Hongjae Lee, Seung-Won JungICCV 2023 · 8 citations
Builds on8
- Channel Attention Is All You Need for Video Frame InterpolationMyungsub Choi, Heewon Kim, Bohyung Han, Ning Xu et al.AAAI 2020 · 362 citations
- XVFI: eXtreme Video Frame InterpolationHyeonjun Sim, Jihyong Oh, Munchurl KimICCV 2021 · 207 citations
- Asymmetric Bilateral Motion Estimation for Video Frame InterpolationJunheum Park, Chul Lee, Chang-Su KimICCV 2021 · 186 citations
- Unsupervised Video Interpolation Using Cycle ConsistencyFitsum A. Reda, Deqing Sun, Aysegul Dundar, Mohammad Shoeybi et al.ICCV 2019 · 93 citations
- Softmax Splatting for Video Frame InterpolationSimon Niklaus, Feng LiuCVPR 2020
Related papers
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
- Disentangled Motion Modeling for Video Frame InterpolationJaihyun Lew, Jooyoung Choi, Chaehun Shin, Dahuin Jung et al.AAAI 2025 · 11 citations
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 213 citations
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel et al.CVPR 2024 · 24 citations
- TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame InterpolationZonglin Lyu, Chen ChenICCV 2025 · 1 citation
