DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
Zheng Chen, Zichen Zou, Kewei Zhang, Xiongfei Su, Xin Yuan, Yong Guo, Yulun Zhang
Abstract
Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training overhead on video data and stringent fidelity demands. To tackle the above issues, we propose DOVE, an efficient one-step diffusion model for real-world VSR. DOVE is obtained by fine-tuning a pretrained video diffusion model (i.e., CogVideoX). To effectively train DOVE, we introduce the latent-pixel training strategy. The strategy employs a two-stage scheme to gradually adapt the model to the video super-resolution task. Meanwhile, we design a video processing pipeline to construct a high-quality dataset tailored for VSR, termed HQ-VSR. Fine-tuning on this dataset further enhances the restoration capability of DOVE. Extensive experiments show that DOVE exhibits comparable or superior performance to multi-step diffusion-based VSR methods. It also offers outstanding inference efficiency, achieving up to a 28 speed-up over existing methods such as MGLD-VSR. Code is available at: https://github.com/zhengchen1999/DOVE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang et al.NeurIPS 2025 · 89 citations
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li et al.CVPR 2026 · 42 citations
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-ResolutionZhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. WongCVPR 2026 · 4 citations
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-ResolutionJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang et al.CVPR 2026 · 1 citation
- PS-SR: Pseudo-Single-Step Video Super-Resolution via Speculative DiffusionAiqiu Wu, Zhaofan Qiu, Ting Yao, Tao MeiCVPR 2026
Builds on36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- TurboVSR: Fantastic Video Upscalers and Where to Find ThemZhongdao Wang, Guodongfang Zhao, Jingjing Ren, Bailan Feng et al.ICCV 2025
- SF-V: Single Forward Video Generation ModelZhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu et al.NeurIPS 2024 · 43 citations
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-ResolutionZiqing Zhang, Kai Liu, Zheng Chen, Xi Li et al.ICML 2026
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image CompressionZheng Chen, Mingde Zhou, Jinpei Guo, Jiale Yuan et al.AAAI 2026 · 1 citation
- SinSR: Diffusion-Based Image Super-Resolution in a Single StepYufei Wang, Wenhan Yang, Xinyuan Chen, Yaohui Wang et al.CVPR 2024 · 110 citations
