Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
Shijun Shi, Jing Xu, Lijing Lu, Zhihang Li, Kai Hu
Abstract
Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised learning and Mamba into pre-trained latent diffusion models. To ensure content consistency across adjacent frames, we enhance the diffusion model with a global spatio-temporal attention mechanism using the Video State-Space block with a 3D Selective Scan module, which reinforces coherence at an affordable computational cost. To further reduce artifacts in generated details, we introduce a self-supervised ControlNet that leverages HR features as guidance and employs contrastive learning to extract degradation-insensitive features from LR videos. Finally, a three-stage training strategy based on a mixture of HR-LR videos is proposed to stabilize VSR training. The proposed Self-supervised ControlNet with Spatio-Temporal Continuous Mamba based VSR algorithm achieves superior perceptual quality than state-of-the-arts on real-world VSR benchmark datasets, validating the effectiveness of the proposed model design and training strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be1834db-52ea-4e56-bdae-b7d0ebf8c57dCited by top-tier papers3
- One-to-All Animation: Alignment-Free Character Animation and Image Pose TransferShijun Shi, Jing Xu, Zhihang Li, Chunli Peng et al.CVPR 2026 · 11 citations
- Compressed-Domain-Aware Online Video Super-ResolutionYuhang Wang, Hai Li, Shujuan Hou, Zhetao Dong et al.CVPR 2026 · 1 citation
- VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent MambaLongmi Gao, Pan GaoCVPR 2026
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- DAM-VSR: Disentanglement of Appearance and Motion for Video Super-ResolutionZhe Kong, Le Li, Yong Zhang, Feng Gao et al.SIGGRAPH 2025 · 6 citations
- One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-ResolutionYujing Sun, Lingchen Sun, Shuaizheng Liu, Rongyuan Wu et al.NeurIPS 2025 · 22 citations
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionShian Du, Menghan Xia, Chang Liu, Xintao Wang et al.CVPR 2025
- Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion ModelCong Cao, Huanjing Yue, Xin Liu, Jingyu YangAAAI 2025 · 7 citations
- Event-based Video Super-Resolution via State Space ModelsZeyu Xiao, Xinchao WangCVPR 2025
