Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
Giannis Daras, Weili Nie, Karsten Kreis, Alex Dimakis, Morteza Mardani, Nikola B. Kovachki, Arash Vahdat
Abstract
Using image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions in the 2D space, and videos as a sequence of continuous warping transformations between different frames. This perspective allows us to train function space diffusion models only on images and utilize them to solve temporally correlated inverse problems. The function space diffusion models need to be equivariant with respect to the underlying spatial transformations. To ensure temporal consistency, we introduce a simple post-hoc test-time guidance towards (self)-equivariant solutions. Our method allows us to deploy state-of-the-art latent diffusion models such as Stable Diffusion XL to solve video inverse problems. We demonstrate the effectiveness of our method for video inpainting and video super-resolution, outperforming existing techniques based on noise transformations. We provide generated video results: https://giannisdaras.github.io/warped_diffusion.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 056fc6aa-abde-40af-97a6-b260b9bea6cdCited by top-tier papers12
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao et al.CVPR 2026 · 3 citations
- Temporal-Consistent Video Restoration with Pre-trained Diffusion ModelsHengkang Wang, Yang Liu, Huidong Liu, Chien-Chih Wang et al.AAAI 2026 · 3 citations
- Generative detail enhancement for physically based materialsSaeed Hadadan, Benedikt Bitterli, Tizian Zeltner, Jan Novák et al.SIGGRAPH 2025 · 3 citations
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim et al.SIGGRAPH 2026
- SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent RepresentationMinho Park, Taewoong Kang, Jooyeol Yun, Sungwon Hwang et al.AAAI 2026
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion ModelsTaesung Kwon, Jong Chul YeICCV 2025
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang et al.ACM MM 2023 · 12 citations
- Solving Video Inverse Problems Using Image Diffusion ModelsTaesung Kwon, Jong Chul YeICLR 2025
- Video Diffusion Models Are Strong Video InpainterMinhyeok Lee, Suhwan Cho, Chajin Shin, Jungho Lee et al.AAAI 2025 · 26 citations
- DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex DegradationsXiaohui Li, Yihao Liu, Shuo Cao, Ziyan Chen et al.ICCV 2025 · 7 citations
