SRFormer: Permuted Self-Attention for Single Image Super-Resolution
Yupeng Zhou, Zhen Li, Chun-Le Guo, Song Bai, Ming-Ming Cheng, Qibin Hou
摘要
Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance but the computation overhead is also considerable. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Our PSA is simple and can be easily applied to existing super-resolution networks based on window self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86dB PSNR score on the Urban100 dataset, which is 0.46dB higher than that of SwinIR but uses fewer parameters and computations. We hope our simple and effective approach can serve as a useful tool for future research in super-resolution model design. Our code is available at https://github.com/HVision-NKU/SRFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- U-DiTs: Downsample Tokens in U-Shaped Diffusion TransformersYuchuan Tian, Zhijun Tu, Hanting Chen, Jie Hu 等NeurIPS 2024 · 被引用 59 次
- See More Details: Efficient Image Super-Resolution by Experts MiningEduard Zamfir, Zongwei Wu, Nancy Mehta, Yulun Zhang 等ICML 2024 · 被引用 38 次
- Image Processing GNN: Breaking Rigidity in Super-ResolutionYuchuan Tian, Hanting Chen, Chao Xu, Yunhe WangCVPR 2024 · 被引用 34 次
- Emulating Self-attention with Convolution for Efficient Image Super-ResolutionDongheon Lee, Seokju Yun, Youngmin RoICCV 2025 · 被引用 19 次
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 被引用 16 次
它引用的顶会 Paper35
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- GRFormer: Grouped Residual Self-Attention for Lightweight Single Image Super-ResolutionYuzhen Li, Zehang Deng, Yuxin Cao, Lihua LiuACM MM 2024 · 被引用 9 次
- N-Gram in Swin Transformers for Efficient Lightweight Image Super-ResolutionHaram Choi, Jeongmin Lee, Jihoon YangCVPR 2023
- Activating More Pixels in Image Super-Resolution TransformerXiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao 等CVPR 2023
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
- MatteFormer: Transformer-Based Image Matting via Prior-TokensGyutae Park, Sungjoon Son, Jaeyoung Yoo, Seho Kim 等CVPR 2022 · 被引用 82 次
