Random Shuffle Transformer for Image Restoration
Jie Xiao, Xueyang Fu, Man Zhou, Hongjian Liu, Zheng-Jun Zha
Abstract
Non-local interactions play a vital role in boosting performance for image restoration. However, local window Transformer has been preferred due to its efficiency for processing high-resolution images. The superiority in efficiency comes at the cost of sacrificing the ability to model nonlocal interactions. In this paper, we present that local window Transformer can also function as modeling non-local interactions. The counterintuitive function is based on the permutationequivariance of self-attention. The basic principle is quite simple: by randomly shuffling the input, local self-attention also has the potential to model non-local interactions without introducing extra parameters. Our random shuffle strategy enjoys elegant theoretical guarantees in extending the local scope. The resulting Transformer dubbed Shuffle-Former is capable of processing high-resolution images efficiently while modeling non-local interactions. Extensive experiments demonstrate the effectiveness of ShuffleFormer across a variety of image restoration tasks, including image denoising, deraining, and deblurring. Code is available at https://github.com/jiexiaou/ ShuffleFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b65a436-29fb-45af-a8ab-471d936759f8Cited by top-tier papers3
- HomoFormer: Homogenized Transformer for Image Shadow RemovalJie Xiao, Xueyang Fu, Yurui Zhu, Dong Li et al.CVPR 2024
- Continuous Adverse Weather Removal via Degradation-Aware DistillationXin Lu, Jie Xiao, Yurui Zhu, Xueyang FuCVPR 2025
- Solving Masked Jigsaw Puzzles with Diffusion Vision TransformersJinyang Liu, Wondmgezahu Teshome, Sandesh Ghimire, Mario Sznaier et al.CVPR 2024
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
Related papers
- Stochastic Window Transformer for Image RestorationJie Xiao, Xueyang Fu, Feng Wu, Zheng-Jun ZhaNeurIPS 2022 · 37 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- KNN Local Attention for Image RestorationHunsang Lee, Hyesong Choi, Kwanghoon Sohn, Dongbo MinCVPR 2022 · 62 citations
- Cross Aggregation Transformer for Image RestorationZheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang et al.NeurIPS 2022 · 274 citations
- LoFormer: Local Frequency Transformer for Image DeblurringXintian Mao, Jiansheng Wang, Xingran Xie, Qingli Li et al.ACM MM 2024 · 44 citations
