ShuffleMixer: An Efficient ConvNet for Image Super-Resolution
Long Sun, Jinshan Pan, Jinhui Tang
Abstract
Lightweight and efficiency are critical drivers for the practical application of image super-resolution (SR) algorithms. We propose a simple and effective approach, ShuffleMixer, for lightweight image super-resolution that explores large convolution and channel split-shuffle operation. In contrast to previous SR models that simply stack multiple small kernel convolutions or complex operators to learn representations, we explore a large kernel ConvNet for mobile-friendly SR design. Specifically, we develop a large depth-wise convolution and two projection layers based on channel splitting and shuffling as the basic component to mix features efficiently. Since the contexts of natural images are strongly locally correlated, using large depth-wise convolutions only is insufficient to reconstruct fine details. To overcome this problem while maintaining the efficiency of the proposed module, we introduce Fused-MBConvs into the proposed network to model the local connectivity of different features. Experimental results demonstrate that the proposed ShuffleMixer is about 6× smaller than the state-of-the-art methods in terms of model parameters and FLOPs while achieving competitive performance. In NTIRE 2022, our primary method won the model complexity track of the Efficient Super-Resolution Challenge [23] . The code is available at https://github.com/sunny2109/MobileSR-NTIRE2022 . Recently, convolutional neural network (CNN) based SR models [8, 9, 1, 16, 25, 45] have achieved impressive reconstruction performance. However, these networks hierarchically extract local features, which highly rely on stacking deeper or more complex models to enlarge the receptive fields for performance improvements. As a result, the required computational budget makes these heavy SR models difficult to deploy on resource-constrained mobile devices in practical applications [44] . To alleviate heavy SR models, various methods have been proposed to reduce model complexity or speed up runtime, including efficient operation design [32, 28, 36, 9, 16, 1, 33, 43, 23, 27] , neural architecture search [6, 35] , knowledge distillation [12, 13] , and structural re-parameterization methodology [7, 23, 44] . These methods are mainly based on improved small spatial convolutions or advanced training strategies, and large kernel convolutions are rarely explored. Moreover, they mostly focus on one of the efficiency indicators and do not perform well in real resource-constrained tasks. Thus, the need to obtain a better trade-off between complexity, latency, and SR quality is imperative. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa9f45c2-26c3-4d33-8099-16d4ccfbf7d1Cited by top-tier papers14
- Spatially-Adaptive Feature Modulation for Efficient Image Super-ResolutionLong Sun, Jiangxin Dong, Jinhui Tang, Jinshan PanICCV 2023 · 211 citations
- Emulating Self-attention with Convolution for Efficient Image Super-ResolutionDongheon Lee, Seokju Yun, Youngmin RoICCV 2025 · 19 citations
- Omnidirectional Image Super-resolution via Bi-projection FusionJiangang Wang, Yuning Cui, Yawen Li, Wenqi Ren et al.AAAI 2024 · 15 citations
- Efficient Single Image Super-Resolution with Entropy Attention and Receptive Field AugmentationXiaole Zhao, Linze Li, Chengxing Xie, Xiaoming Zhang et al.ACM MM 2024 · 12 citations
- Unveiling Details in the Dark: Simultaneous Brightening and Zooming for Low-Light Image EnhancementZiyu Yue, Jiaxin Gao, Zhixun SuAAAI 2024 · 10 citations
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- Self-feature Learning: An Efficient Deep Lightweight Network for Image Super-resolutionJun Xiao, Qian Ye, Rui Zhao, Kin-Man Lam et al.ACM MM 2021 · 17 citations
- CAMixerSR: Only Details Need More "Attention"Yan Wang, Yi Liu, Shijie Zhao, Junlin Li et al.CVPR 2024 · 60 citations
- Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolutionBin Sun, Yulun Zhang, Songyao Jiang, Yun FuAAAI 2023 · 72 citations
- SplitSR: An End-to-End Approach to Super-Resolution on Mobile DevicesXin Liu, Yuang Li, Josh Fromm, Yuntao Wang et al.UbiComp 2021 · 29 citations
- Edge-oriented Convolution Block for Real-time Super Resolution on Mobile DevicesXindong Zhang, Hui Zeng, Lei ZhangACM MM 2021 · 229 citations
