PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
Shian Du, Menghan Xia, Chang Liu, Xintao Wang, Jing Wang, Pengfei Wan, Di Zhang, Xiangyang Ji
Abstract
Pre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we make the first exploration into utilizing video diffusion priors for patch-wise VSR. This is non-trivial because pre-trained video diffusion models are not native for patch-level detail generation. To mitigate this challenge, we propose an innovative approach, called PatchVSR, which integrates a dual-stream adapter for conditional guidance. The patch branch extracts features from input patches to maintain content fidelity while the global branch extracts context features from the resized full video to bridge the generation gap caused by incomplete semantics of patches. Particularly, we also inject the patch's location information into the model to better contextualize patch synthesis within the global video frame. Experiments demonstrate that our method can synthesize high-fidelity, high-resolution details at the patch level. A tailor-made multi-patch joint modulation is proposed to ensure visual consistency across individually enhanced patches. Due to the flexibility of our patch-based paradigm, we can achieve highly competitive 4K VSR based on a 512×512 resolution base model, with extremely high efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a9d8aa2-bf77-4536-bb58-8f56cb0024a6Cited by top-tier papers2
- Compressed-Domain-Aware Online Video Super-ResolutionYuhang Wang, Hai Li, Shujuan Hou, Zhetao Dong et al.CVPR 2026 · 1 citation
- Seeing the Unseen: Zooming in the Dark with Event CamerasDachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang et al.AAAI 2026
Builds on32
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-ResolutionZiqing Zhang, Kai Liu, Zheng Chen, Xi Li et al.ICML 2026
- WEVSR: Video Diffusion Generators for Real-World Video Super‑Resolution with Wavelet-Enhanced VAE EncoderYuying Chen, Liu, Linyan Jiang, Qifan Gao et al.ICML 2026
- One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-ResolutionYujing Sun, Lingchen Sun, Shuaizheng Liu, Rongyuan Wu et al.NeurIPS 2025 · 22 citations
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterJianhui Zhang, Sheng Cheng, Qirui Sun, Jia Liu et al.ICCV 2025 · 1 citation
- DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion TransformerQingji Dong, Hang Dong, Mingqin Chen, Rui Zhang et al.CVPR 2026 · 1 citation
