STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
Junyang Chen, Jiangxin Dong, Long Sun, Yixin Yang, Jinshan Pan
Abstract
We present STCDiT, a video super-resolution framework built upon a pre-trained video diffusion model, aiming to restore structurally faithful and temporally stable videos from degraded inputs, even under complex camera motions. The main challenges lie in maintaining temporal stability during reconstruction and preserving structural fidelity during generation. To address these challenges, we first develop a motion-aware VAE reconstruction method that performs segment-wise reconstruction, with each segment clip exhibiting uniform motion characteristic, thereby effectively handling videos with complex camera motions. Moreover, we observe that the first-frame latent extracted by the VAE encoder in each clip, termed the anchor-frame latent, remains unaffected by temporal compression and retains richer spatial structural information than subsequent frame latents. We further develop an anchor-frame guidance approach that leverages structural information from anchor frames to constrain the generation process and improve structural fidelity of video features. Coupling these two designs enables the video diffusion model to achieve high-quality video super-resolution. Extensive experiments show that STCDiT outperforms state-of-the-art methods in terms of structural fidelity and temporal consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b4a5c96-e980-42e3-963f-2d60870ef2c1Builds on29
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen et al.ICCV 2023 · 371 citations
Related papers
- DAM-VSR: Disentanglement of Appearance and Motion for Video Super-ResolutionZhe Kong, Le Li, Yong Zhang, Feng Gao et al.SIGGRAPH 2025 · 6 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
- VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAEYazhou Xing, Yang Fei, Yingqing He, Jingye Chen et al.ICCV 2025 · 2 citations
- Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-ResolutionZhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao et al.CVPR 2024
- AnchorSync: Global Consistency Optimization for Long Video EditingZichi Liu, Yinggui Wang, Tao Wei, Chao MaACM MM 2025
