Cross Paradigm Representation and Alignment Transformer for Image Deraining
Shun Zou, Yi Zou, Juncheng Li, Guangwei Gao, Guo-Jun Qi
摘要
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
- FFA-Net: Feature Fusion Attention Network for Single Image DehazingXu Qin, Zhilin Wang, Yuanchao Bai, Xiaodong Xie 等AAAI 2020 · 被引用 1,828 次
相关 Paper
- Learning A Sparse Transformer Network for Effective Image DerainingXiang Chen, Hao Li, Mingqiang Li, Jinshan PanCVPR 2023
- Magic ELF: Image Deraining Meets Association Learning and TransformerKui Jiang, Zhongyuan Wang, Chen Chen, Zheng Wang 等ACM MM 2022 · 被引用 99 次
- Rethinking Multi-Scale Representations in Deep Deraining TransformerHongming Chen, Xiang Chen, Jiyang Lu, Yufeng LiAAAI 2024 · 被引用 46 次
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain StreaksSixiang Chen, Tian Ye, Jinbin Bai, Erkang Chen 等ICCV 2023 · 被引用 72 次
