FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super Resolution
Junhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li, Yihao Liu, Chun Yuan, Tianfan Xue
摘要
Diffusion models have recently advanced video restoration, but applying them to real-world video super-resolution (VSR) remains challenging due to high latency, prohibitive computation, and poor generalization to ultra-high resolutions. Our goal in this work is to make diffusion-based VSR practical by achieving efficiency, scalability, and real-time performance. To this end, we propose FlashVSR, the first diffusion-based one-step streaming framework towards realtime VSR. FlashVSR runs at ∼17 FPS for 768 × 1408 videos on a single A100 GPU by combining three complementary innovations: (i) a train-friendly threestage distillation pipeline that enables streaming super-resolution, (ii) localityconstrained sparse attention that cuts redundant computation while bridging the train-test resolution gap, and (iii) a tiny conditional decoder that accelerates reconstruction without sacrificing quality. To support large-scale training, we also construct VSR-120K, a new dataset with 120k videos and 180k images. Extensive experiments show that FlashVSR scales reliably to ultra-high resolutions and achieves state-of-the-art performance with up to ∼ 12× speedup over prior one-step diffusion VSR models. We will release the code, pretrained models, and dataset to foster future research in efficient diffusion-based VSR at https://zhuang2002.github.io/FlashVSR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 等ICML 2026 · 被引用 24 次
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang 等CVPR 2026 · 被引用 10 次
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-ResolutionZhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. WongCVPR 2026 · 被引用 4 次
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
- SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian SplattingXiang Feng, Xiangbo Wang, Tieshi Zhong, Chengkai Wang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper41
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
相关 Paper
- UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion SpaceYong Liu, Jinshan Pan, Yinchuan Li, Qingji Dong 等ACM MM 2025 · 被引用 3 次
- TurboVSR: Fantastic Video Upscalers and Where to Find ThemZhongdao Wang, Guodongfang Zhao, Jingjing Ren, Bailan Feng 等ICCV 2025
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with TransformersMinguk Kang, Suha KwakCVPR 2026 · 被引用 1 次
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-ResolutionZiqing Zhang, Kai Liu, Zheng Chen, Xi Li 等ICML 2026
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionShian Du, Menghan Xia, Chang Liu, Xintao Wang 等CVPR 2025
