Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
Giannis Daras, Weili Nie, Karsten Kreis, Alex Dimakis, Morteza Mardani, Nikola B. Kovachki, Arash Vahdat
摘要
Using image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions in the 2D space, and videos as a sequence of continuous warping transformations between different frames. This perspective allows us to train function space diffusion models only on images and utilize them to solve temporally correlated inverse problems. The function space diffusion models need to be equivariant with respect to the underlying spatial transformations. To ensure temporal consistency, we introduce a simple post-hoc test-time guidance towards (self)-equivariant solutions. Our method allows us to deploy state-of-the-art latent diffusion models such as Stable Diffusion XL to solve video inverse problems. We demonstrate the effectiveness of our method for video inpainting and video super-resolution, outperforming existing techniques based on noise transformations. We provide generated video results: https://giannisdaras.github.io/warped_diffusion.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
- Temporal-Consistent Video Restoration with Pre-trained Diffusion ModelsHengkang Wang, Yang Liu, Huidong Liu, Chien-Chih Wang 等AAAI 2026 · 被引用 3 次
- Generative detail enhancement for physically based materialsSaeed Hadadan, Benedikt Bitterli, Tizian Zeltner, Jan Novák 等SIGGRAPH 2025 · 被引用 3 次
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim 等SIGGRAPH 2026
- SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent RepresentationMinho Park, Taewoong Kang, Jooyeol Yun, Sungwon Hwang 等AAAI 2026
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion ModelsTaesung Kwon, Jong Chul YeICCV 2025
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang 等ACM MM 2023 · 被引用 12 次
- Solving Video Inverse Problems Using Image Diffusion ModelsTaesung Kwon, Jong Chul YeICLR 2025
- Video Diffusion Models Are Strong Video InpainterMinhyeok Lee, Suhwan Cho, Chajin Shin, Jungho Lee 等AAAI 2025 · 被引用 26 次
- DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex DegradationsXiaohui Li, Yihao Liu, Shuo Cao, Ziyan Chen 等ICCV 2025 · 被引用 7 次
