CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
Xin Liu, Jie Liu, Jie Tang, Gangshan Wu
Abstract
Transformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution images into local windows, axial stripes, or dilated windows. SR typically leverages the redundancy of images for reconstruction, and this redundancy appears not only in local regions but also in long-range regions. However, these methods limit attention computation to contentagnostic local regions, limiting directly the ability of attention to capture long-range dependency. To address these issues, we propose a lightweight Content-Aware Token Aggregation Network (CATANet). Specifically, we propose an efficient Content-Aware Token Aggregation module for aggregating long-range content-similar tokens, which shares token centers across all image tokens and updates them only during the training phase. Then we utilize intra-group self-attention to enable long-range information interaction. Moreover, we design an inter-group cross-attention to further enhance global information interaction. The experimental results show that, compared with the state-of-theart cluster-based method SPIN, our method achieves superior performance, with a maximum PSNR improvement of 0.33dB and nearly double the inference speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 333bd177-4e35-41d4-a723-cd014ba57b2bCited by top-tier papers2
- Compressed-Domain-Aware Online Video Super-ResolutionYuhang Wang, Hai Li, Shujuan Hou, Zhetao Dong et al.CVPR 2026 · 1 citation
- IAFMNet: Information-Aware Feature Modulation for Efficient Super-ResolutionJunwei Xu, Mengzu Liu, Zhenyu Wang, Fangfang Wu et al.CVPR 2026
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
Related papers
- Lightweight Image Super-Resolution with Superpixel Token InteractionAiping Zhang, Wenqi Ren, Yi Liu, Xiaochun CaoICCV 2023 · 59 citations
- Cross Aggregation Transformer for Image RestorationZheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang et al.NeurIPS 2022 · 274 citations
- From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolutionJie Liu, Chao Chen, Jie Tang, Gangshan WuAAAI 2023 · 26 citations
- Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-ResolutionKaram Park, Jae Woong Soh, Nam Ik ChoAAAI 2025 · 20 citations
- DLGSANet: Lightweight Dynamic Local and Global Self-Attention Network for Image Super-ResolutionXiang Li, Jiangxin Dong, Jinhui Tang, Jinshan PanICCV 2023 · 69 citations
