CFAT: Unleashing Triangular Windows for Image Super-resolution
Abhisek Ray, Gaurav Kumar, Maheshkumar H. Kolekar
Abstract
Transformer-based models have revolutionized the field of image super-resolution (SR) by harnessing their inherent ability to capture complex contextual features. The overlapping rectangular shifted window technique used in transformer architecture nowadays is a common practice in super-resolution models to improve the quality and robustness of image upscaling. However, it suffers from distortion at the boundaries and has limited unique shifting modes. To overcome these weaknesses, we propose a non-overlapping triangular window technique that synchronously works with the rectangular one to mitigate boundary-level distortion and allows the model to access more unique sifting modes. In this paper, we propose a Composite Fusion Attention Transformer (CFAT) that incorporates triangular-rectangular window-based local attention with a channel-based global attention technique in image super-resolution. As a result, CFAT enables attention mechanisms to be activated on more image pixels and captures long-range, multi-scale features to improve SR performance. The extensive experimental results and ablation study demonstrate the effectiveness of CFAT in the SR domain. Our proposed model shows a significant 0.7 dB performance improvement over other state-of-the-art SR architectures. Find the code Here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c22c4fc0-138d-4bd3-9d8f-49e0ae1e9c6fCited by top-tier papers1
Ask how each one uses itBuilds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
Related papers
- Activating More Pixels in Image Super-Resolution TransformerXiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao et al.CVPR 2023
- Cross Aggregation Transformer for Image RestorationZheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang et al.NeurIPS 2022 · 274 citations
- Recursive Generalization Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong et al.ICLR 2024 · 81 citations
- Dual Aggregation Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong et al.ICCV 2023 · 345 citations
- Emulating Self-attention with Convolution for Efficient Image Super-ResolutionDongheon Lee, Seokju Yun, Youngmin RoICCV 2025 · 19 citations
