LDMVFI: Video Frame Interpolation with Latent Diffusion Models
Duolikun Danier, Fan Zhang, David Bull
Abstract
Existing works on video frame interpolation (VFI) mostly employ deep neural networks that are trained by minimizing the L1, L2, or deep feature space distance (e.g. VGG loss) between their outputs and ground-truth frames. However, recent works have shown that these metrics are poor indicators of perceptual VFI quality. Towards developing perceptually-oriented VFI methods, in this work we propose latent diffusion model-based VFI, LDMVFI. This approaches the VFI problem from a generative perspective by formulating it as a conditional generation problem. As the first effort to address VFI using latent diffusion models, we rigorously benchmark our method on common test sets used in the existing VFI literature. Our quantitative experiments and user study indicate that LDMVFI is able to interpolate video content with favorable perceptual quality compared to the state of the art, even in the high-resolution regime. Our code is available at https://github.com/danier97/LDMVFI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54f23846-9ec6-41ad-9561-3376bbcd1975Cited by top-tier papers39
- Norm-guided latent space exploration for text-to-image generationDvir Samuel, Rami Ben-Ari, Nir Darshan, Haggai Maron et al.NeurIPS 2023 · 49 citations
- Perception-Oriented Video Frame Interpolation via Asymmetric BlendingGuangyang Wu, Xin Tao, Changlin Li, Wenyi Wang et al.CVPR 2024 · 16 citations
- Disentangled Motion Modeling for Video Frame InterpolationJaihyun Lew, Jooyoung Choi, Chaehun Shin, Dahuin Jung et al.AAAI 2025 · 11 citations
- Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Yijie Yu, Ling Yang, Chujun Qin et al.ACM MM 2024 · 10 citations
- High-Resolution Frame Interpolation with Patch-based Cascaded DiffusionJunhwa Hur, Charles Herrmann, Saurabh Saxena, Janne Kontkanen et al.AAAI 2025 · 8 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Frame Interpolation with Consecutive Brownian Bridge DiffusionZonglin Lyu, Ming Li, Jianbo Jiao, Chen ChenACM MM 2024 · 7 citations
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
- Realtime Video Frame Interpolation using One-Step Diffusion SamplingYongrui Ma, Shijie Zhao, Mingde Yao, Junlin Li et al.ICLR 2026
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion TransformersXinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng et al.CVPR 2026 · 4 citations
- TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame InterpolationZonglin Lyu, Chen ChenICCV 2025 · 1 citation
