RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement
Gang He, Weiran Wang, Guancheng Quan, Shihao Wang, Dajiang Zhou, Yunsong Li
Abstract
Quality degradation from video compression manifests both spatially along texture edges and temporally with continuous motion changes. Despite recent advances, extracting aligned spatiotemporal information from adjacent frames remains challenging. This is mainly due to limitations in receptive field size and computational complexity, which makes existing methods struggle to efficiently enhance video quality. To address this issue, we propose RivuletMLP, an MLP-based network architecture. Specifically, our framework first employs a Dynamically Guided Deformable Alignment (DDA) module to adaptively explore and align multi-frame feature information. Subsequently, we introduce two modules for feature reconstruction: a Spatiotemporal Feature Flow (SFF) Module and a Benign Selection Compensation (BSC) module. The SFF module establishes non-local dependencies through an innovative feature permutation mechanism. Additionally, the BSC module utilizes a collaborative strategy of deep feature extraction and local region refinement to alleviate inter frame motion discontinuity caused by compression. Experimental results demonstrate that RivuletMLP achieves superior computational efficiency while maintaining powerful reconstruction capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8d1ff21-710a-49c5-aa71-8f547e3225f6Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- MAXIM: Multi-Axis MLP for Image ProcessingZhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang et al.CVPR 2022 · 550 citations
Related papers
- Spatio-Temporal Deformable Convolution for Compressed Video Quality EnhancementJianing Deng, Li Wang, Shiliang Pu, Cheng ZhuoAAAI 2020 · 168 citations
- FVC: A New Framework Towards Deep Video Compression in Feature SpaceZhihao Hu, Guo Lu, Dong XuCVPR 2021
- Non-Local ConvLSTM for Video Compression Artifact ReductionYi Xu, Longwen Gao, Kai Tian, Shuigeng Zhou et al.ICCV 2019 · 70 citations
- Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact ReductionMinyi Zhao, Yi Xu, Shuigeng ZhouACM MM 2021 · 61 citations
- TDAN: Temporally-Deformable Alignment Network for Video Super-ResolutionYapeng Tian, Yulun Zhang, Yun Fu, Chenliang XuCVPR 2020
