SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
Qi Tang, Yao Zhao, Meiqin Liu, Chao Yao
Abstract
Diffusion-based Video Super-Resolution (VSR) is renowned for generating perceptually realistic videos, yet it grapples with maintaining detail consistency across frames due to stochastic fluctuations. The traditional approach of pixel-level alignment is ineffective for diffusion-processed frames because of iterative disruptions. To overcome this, we introduce SeeClear--a novel VSR framework leveraging conditional video generation, orchestrated by instance-centric and channel-wise semantic controls. This framework integrates a Semantic Distiller and a Pixel Condenser, which synergize to extract and upscale semantic details from low-resolution frames. The Instance-Centric Alignment Module (InCAM) utilizes video-clip-wise tokens to dynamically relate pixels within and across frames, enhancing coherency. Additionally, the Channel-wise Texture Aggregation Memory (CaTeGory) infuses extrinsic knowledge, capitalizing on long-standing semantic textures. Our method also innovates the blurring diffusion process with the ResShift mechanism, finely balancing between sharpness and diffusion effects. Comprehensive experiments confirm our framework's advantage over state-of-the-art diffusion-based VSR techniques. The code is available: https://github.com/Tang1705/SeeClear-NeurIPS24.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a977b6fb-c8d3-456d-8048-d8102f2899c5Cited by top-tier papers4
- DAM-VSR: Disentanglement of Appearance and Motion for Video Super-ResolutionZhe Kong, Le Li, Yong Zhang, Feng Gao et al.SIGGRAPH 2025 · 6 citations
- QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality AssessmentGuohua Zhang, Jian Jin, Meiqin Liu, Chao Yao et al.CVPR 2026 · 3 citations
- Spatial Imputation Drives Cross-Domain Alignment for EEG ClassificationHongjun Liu, Chao Yao, Yalan Zhang, Xiaokun Wang et al.ACM MM 2025 · 1 citation
- Proper Hölder-Kullback Dirichlet Diffusion: A Framework for High Dimensional Generative ModelingWanpeng Zhang, Yuhao Fang, Xihang Qiu, Jiarong Cheng et al.NeurIPS 2025
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 958 citations
- ResShift: Efficient Diffusion Model for Image Super-resolution by Residual ShiftingZongsheng Yue, Jianyi Wang, Chen Change LoyNeurIPS 2023 · 646 citations
Related papers
- Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolutionQi Tang, Yao Zhao, Meiqin Liu, Jian Jin et al.AAAI 2024 · 10 citations
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionShian Du, Menghan Xia, Chang Liu, Xintao Wang et al.CVPR 2025
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li et al.CVPR 2026 · 42 citations
- CoSeR: Bridging Image and Language for Cognitive Super-ResolutionHaoze Sun, Wenbo Li, Jianzhuang Liu, Haoyu Chen et al.CVPR 2024 · 43 citations
- Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned FeaturesJingyi Xu, Meisong Zheng, Ying Chen, Minglang Qiao et al.CVPR 2026 · 1 citation
