SimpleGVR: A Simple Baseline for Latent-Cascaded Generative Video Super-Resolution
Liangbin Xie, Yu Li, Shian Du, Menghan Xia, Xintao Wang, Fanghua Yu, Ziyan Chen, Pengfei Wan, Jiantao Zhou, Chao Dong
Abstract
Cascaded pipelines, which use a base text-to-video (T2V) model for low-resolution content and a video super-resolution (VSR) model for high-resolution details, are a prevailing strategy for efficient video synthesis. However, current works suffer from two key limitations: an inefficient pixel-space interface that introduces non-trivial computational overhead, and mismatched degradation strategies that compromise the visual quality of AIGC content. To address these issues, we introduce SimpleGVR, a lightweight VSR model designed to operate entirely within the latent space. Key to SimpleGVR are a latent upsampler for effective, detail-preserving conditioning of the high-resolution synthesis, and two degradation strategies (flow-based and model-guided) to ensure better alignment with the upstream T2V model. To further enhance the performance and practical applicability of SimpleGVR, we introduce a set of crucial training optimizations: a detail-aware timestep sampler, a suitable noise augmentation range, and an efficient interleaving temporal unit mechanism for long-video handling. Extensive experiments demonstrate the superiority of our framework over existing methods, with ablation studies confirming the efficacy of each design. Our work establishes a simple yet effective baseline for cascaded video super-resolution generation, offering practical insights to guide future advancements in efficient cascaded systems. Video visual comparisons are available at https://simplegvr.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 177a022c-aef9-4393-8a75-5a32fa65173eCited by top-tier papers2
- CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video GenerationKaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai et al.CVPR 2026 · 7 citations
- WEVSR: Video Diffusion Generators for Real-World Video Super‑Resolution with Wavelet-Enhanced VAE EncoderYuying Chen, Liu, Linyan Jiang, Qifan Gao et al.ICML 2026
Builds on28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- TurboVSR: Fantastic Video Upscalers and Where to Find ThemZhongdao Wang, Guodongfang Zhao, Jingjing Ren, Bailan Feng et al.ICCV 2025
- DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex DegradationsXiaohui Li, Yihao Liu, Shuo Cao, Ziyan Chen et al.ICCV 2025 · 7 citations
- Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionRui Xie, Yinhong Liu, Penghao Zhou, Chen Zhao et al.ICCV 2025 · 11 citations
- VideoGigaGAN: Towards Detail-rich Video Super-ResolutionYiran Xu, Taesung Park, Richard Zhang, Yang Zhou et al.CVPR 2025
- LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency ExpertsChen Zhao, Jiawei Chen, Hongyu Li, Zhuoliang Kang et al.ICML 2026 · 16 citations
