Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
Shijun Shi, Jing Xu, Lijing Lu, Zhihang Li, Kai Hu
摘要
Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised learning and Mamba into pre-trained latent diffusion models. To ensure content consistency across adjacent frames, we enhance the diffusion model with a global spatio-temporal attention mechanism using the Video State-Space block with a 3D Selective Scan module, which reinforces coherence at an affordable computational cost. To further reduce artifacts in generated details, we introduce a self-supervised ControlNet that leverages HR features as guidance and employs contrastive learning to extract degradation-insensitive features from LR videos. Finally, a three-stage training strategy based on a mixture of HR-LR videos is proposed to stabilize VSR training. The proposed Self-supervised ControlNet with Spatio-Temporal Continuous Mamba based VSR algorithm achieves superior perceptual quality than state-of-the-arts on real-world VSR benchmark datasets, validating the effectiveness of the proposed model design and training strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- One-to-All Animation: Alignment-Free Character Animation and Image Pose TransferShijun Shi, Jing Xu, Zhihang Li, Chunli Peng 等CVPR 2026 · 被引用 11 次
- Compressed-Domain-Aware Online Video Super-ResolutionYuhang Wang, Hai Li, Shujuan Hou, Zhetao Dong 等CVPR 2026 · 被引用 1 次
- VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent MambaLongmi Gao, Pan GaoCVPR 2026
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- DAM-VSR: Disentanglement of Appearance and Motion for Video Super-ResolutionZhe Kong, Le Li, Yong Zhang, Feng Gao 等SIGGRAPH 2025 · 被引用 6 次
- One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-ResolutionYujing Sun, Lingchen Sun, Shuaizheng Liu, Rongyuan Wu 等NeurIPS 2025 · 被引用 22 次
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionShian Du, Menghan Xia, Chang Liu, Xintao Wang 等CVPR 2025
- Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion ModelCong Cao, Huanjing Yue, Xin Liu, Jingyu YangAAAI 2025 · 被引用 7 次
- Event-based Video Super-Resolution via State Space ModelsZeyu Xiao, Xinchao WangCVPR 2025
