High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
Junhwa Hur, Charles Herrmann, Saurabh Saxena, Janne Kontkanen, Wei-Sheng Lai, Yichang Shih, Michael Rubinstein, David J. Fleet, Deqing Sun
Abstract
Despite the recent progress, existing frame interpolation methods still struggle with processing extremely high resolution input and handling challenging cases such as repetitive textures, thin objects, and large motion. To address these issues, we introduce a patch-based cascaded pixel diffusion model for high resolution frame interpolation, HiFI, that excels in these scenarios while achieving competitive performance on standard benchmarks. Cascades, which generate a series of images from low to high resolution, can help significantly with large or complex motion that require both global context for a coarse solution and detailed context for high resolution output. However, contrary to prior work on cascaded diffusion models which perform diffusion on increasingly large resolutions, we use a single model that always performs diffusion at the same resolution and upsamples by processing patches of the inputs and the prior solution. At inference time, this drastically reduces memory usage and allows a single model, solving both frame interpolation (base model's task) and spatial up-sampling, saving training cost as well. HiFI excels at high-resolution images and complex repeated textures that require global context, achieving comparable or state-of-the-art performance on various benchmarks (Vimeo, Xiph, X-Test, and SEPE-8K). We further introduce a new dataset, LaMoR, that focuses on particularly challenging cases, and HiFI significantly outperforms other baselines. Please visit our project page for video results: https://hifi-diffusion.github.io * These authors contributed equally. † DF is also affiliated with the University of Toronto and the Vector Institute.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d77bc126-bb77-4e88-898d-bdea99caa2d5Cited by top-tier papers3
- Imagine How To Change: Explicit Procedure Modeling for Change CaptioningJiayang Sun, Zixin Guo, Min Cao, Guibo Zhu et al.ICLR 2026 · 1 citation
- Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large MotionZaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen et al.NeurIPS 2025
- FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded DenoisingHaoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen et al.CVPR 2026
Builds on35
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Label-Efficient Semantic Segmentation with Diffusion ModelsDmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov et al.ICLR 2022 · 700 citations
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
Related papers
- Hierarchical Patch Diffusion Models for High-Resolution Video GenerationIvan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Sergey TulyakovCVPR 2024 · 14 citations
- Video Interpolation with Diffusion ModelsSiddhant Jain, Daniel Watson, Eric Tabellion, Aleksander Holynski et al.CVPR 2024
- simple diffusion: End-to-end diffusion for high resolution imagesEmiel Hoogeboom, Jonathan Heek, Tim SalimansICML 2023 · 403 citations
- Patched Denoising Diffusion Models For High-Resolution Image SynthesisZheng Ding, Mengqi Zhang, Jiajun Wu, Zhuowen TuICLR 2024 · 55 citations
- Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer NormalizationQihao Liu, Zhanpeng Zeng, Ju He, Qihang Yu et al.NeurIPS 2024
