Unfolding Once is Enough: A Deployment-Friendly Transformer Unit for Super-Resolution
Yong Liu, Hang Dong, Boyang Liang, Songwei Liu, Qingji Dong, Kai Chen, Fangmin Chen, Lean Fu, Fei Wang
Abstract
Recent years have witnessed a few attempts of vision transformers for single image super-resolution (SISR). Since the high resolution of intermediate features in SISR models increases memory and computational requirements, efficient SISR transformers are more favored. Based on some popular transformer backbone, many methods have explored reasonable schemes to reduce the computational complexity of the self-attention module while achieving impressive performance. However, these methods only focus on the performance on the training platform (e.g., Pytorch/Tensorflow) without further optimization for the deployment platform (e.g., TensorRT). Therefore, they inevitably contain some redundant operators, posing challenges for subsequent deployment in real-world applications. In this paper, we propose a deployment-friendly transformer unit, namely UFONE (i.e., UnFolding ONce is Enough), to alleviate these problems. In each UFONE, we introduce an Inner-patch Transformer Layer (ITL) to efficiently reconstruct the local structural information from patches and a Spatial-Aware Layer (SAL) to exploit the long-range dependencies between patches. Based on UFONE, we propose a Deployment-friendly Inner-patch Transformer Network (DITN) for the SISR task, which can achieve favorable performance with low latency and memory usage on both training and deployment platforms. Furthermore, to further boost the deployment efficiency of the proposed DITN on TensorRT, we also provide an efficient substitution for layer normalization and propose a fusion optimization strategy for specific operators. Extensive experiments show that our models can achieve competitive results in terms of qualitative and quantitative performance with high deployment efficiency. Code is available at https://github.com/yongliuy/DITN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Emulating Self-attention with Convolution for Efficient Image Super-ResolutionDongheon Lee, Seokju Yun, Youngmin RoICCV 2025 · 19 citations
- PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-ResolutionYong Liu, Hang Dong, Jinshan Pan, Qingji Dong et al.ICCV 2025 · 2 citations
- DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion TransformerQingji Dong, Hang Dong, Mingqin Chen, Rui Zhang et al.CVPR 2026 · 1 citation
- CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionXin Liu, Jie Liu, Jie Tang, Gangshan WuCVPR 2025
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
Related papers
- SHViT: Single-Head Vision Transformer with Memory Efficient Macro DesignSeokju Yun, Youngmin RoCVPR 2024 · 117 citations
- Spatially-Adaptive Feature Modulation for Efficient Image Super-ResolutionLong Sun, Jiangxin Dong, Jinhui Tang, Jinshan PanICCV 2023 · 211 citations
- Effective Diffusion Transformer Architecture for Image Super-ResolutionKun Cheng, Lei Yu, Zhijun Tu, Xiao He et al.AAAI 2025 · 26 citations
- Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-ResolutionKaram Park, Jae Woong Soh, Nam Ik ChoAAAI 2025 · 20 citations
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsYang Bai, Wenqian Zhao, Shuo Yin, Zixiao Wang et al.EMNLP 2023 · 2 citations
