Video Compression Artifact Reduction by Fusing Motion Compensation and Global Context in a Swin-CNN Based Parallel Architecture
Xinjian Zhang, Su Yang, Wuyang Luo, Longwen Gao, Weishan Zhang
Abstract
Video Compression Artifact Reduction aims to reduce the artifacts caused by video compression algorithms and improve the quality of compressed video frames. The critical challenge in this task is to make use of the redundant high-quality information in compressed frames for compensation as much as possible. Two important possible compensations: Motion compensation and global context, are not comprehensively considered in previous works, leading to inferior results. The key idea of this paper is to fuse the motion compensation and global context together to gain more compensation information to improve the quality of compressed videos. Here, we propose a novel Spatio-Temporal Compensation Fusion (STCF) framework with the Parallel Swin-CNN Fusion (PSCF) block, which can simultaneously learn and merge the motion compensation and global context to reduce the video compression artifacts. Specifically, a temporal self-attention strategy based on shifted windows is developed to capture the global context in an efficient way, for which we use the Swin transformer layer in the PSCF block. Moreover, an additional Ada-CNN layer is applied in the PSCF block to extract the motion compensation. Experimental results demonstrate that our proposed STCF framework outperforms the state-of-the-art methods up to 0.23dB (27% improvement) on the MFQEv2 dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8636a378-9c6d-4f40-b3f8-c5f1eb6e0abbCited by top-tier papers2
- RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality EnhancementGang He, Weiran Wang, Guancheng Quan, Shihao Wang et al.CVPR 2025
- Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality EnhancementLiuhan Peng, Shuai Li, Yanbo Gao, Mao Ye et al.AAAI 2026
Builds on6
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Spatio-Temporal Deformable Convolution for Compressed Video Quality EnhancementJianing Deng, Li Wang, Shiliang Pu, Cheng ZhuoAAAI 2020 · 168 citations
- Non-Local ConvLSTM for Video Compression Artifact ReductionYi Xu, Longwen Gao, Kai Tian, Shuigeng Zhou et al.ICCV 2019 · 70 citations
Related papers
- The Devil Is in the Details: Window-based Attention for Image CompressionRenjie Zou, Chunfeng Song, Zhaoxiang ZhangCVPR 2022 · 260 citations
- Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact ReductionMinyi Zhao, Yi Xu, Shuigeng ZhouACM MM 2021 · 61 citations
- Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersZhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang et al.ACM MM 2023 · 20 citations
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
- Transcoded Video Restoration by Temporal Spatial Auxiliary NetworkLi Xu, Gang He, Jinjia Zhou, Jie Lei et al.AAAI 2022 · 17 citations
