Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution
Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo, Chen Change Loy
摘要
Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and quality. Extensive experiments show that Upscale-A-Video surpasses existing methods in both synthetic and real-world benchmarks, as well as in AI-generated videos, showcasing impressive visual realism and temporal consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion PriorsXi Liu, Chaoyi Zhou, Siyu HuangNeurIPS 2024 · 被引用 127 次
- SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-TrainingJianyi Wang, Shanchuan Lin, Zhijie Lin, Yuxi Ren 等ICLR 2026 · 被引用 51 次
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li 等CVPR 2026 · 被引用 42 次
- Recammaster: Camera-Controlled Generative Rendering From a Single VideoJianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang 等ICCV 2025 · 被引用 33 次
- FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationShilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge 等AAAI 2026 · 被引用 30 次
它引用的顶会 Paper48
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 被引用 219 次
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li 等ACM MM 2025 · 被引用 2 次
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-ResolutionJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang 等CVPR 2026 · 被引用 1 次
- Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionRui Xie, Yinhong Liu, Penghao Zhou, Chen Zhao 等ICCV 2025 · 被引用 11 次
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 等ICLR 2024 · 被引用 175 次
