Emulating Self-attention with Convolution for Efficient Image Super-Resolution
Dongheon Lee, Seokju Yun, Youngmin Ro
摘要
In this paper, we tackle the high computational overhead of Transformers for efficient image super-resolution (SR). Motivated by the observations of self-attention's inter-layer repetition, we introduce a convolutionized self-attention module named Convolutional Attention (ConvAttn) that emulates self-attention's long-range modeling capability and instance-dependent weighting with a single shared large kernel and dynamic kernels. By utilizing the ConvAttn module, we significantly reduce the reliance on self-attention and its involved memory-bound operations while maintaining the representational capability of Transformers. Furthermore, we overcome the challenge of integrating flash attention into the lightweight SR regime, effectively mitigating self-attention's inherent memory bottleneck. We scale up the window size to with flash attention rather than proposing an intricate self-attention module, significantly improving PSNR by 0.31 dB on Urban while reducing latency and memory usage by and . Building on these approaches, our proposed network, termed Emulating Self-attention with Convolution (ESC), notably improves PSNR by 0.27 dB on Urban compared to HiT-SRF, reducing the latency and memory usage by and , respectively. Extensive experiments demonstrate that our ESC maintains the ability for long-range modeling, data scalability, and the representational power of Transformers despite most self-attention being replaced by the ConvAttn module.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LRDUN: A Low-Rank Deep Unfolding Network for Efficient Spectral Compressive ImagingHE HUANG, Yujun Guo, Wei HeCVPR 2026 · 被引用 3 次
- UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-ResolutionCao Thien Tan, Phan Thi Thu Trang, Do Nghiem Duc, Ho Ngoc Anh 等CVPR 2026 · 被引用 2 次
- Joint Geometric and Trajectory Consistency Learning for One-Step Real-World Super-ResolutionChengyan Deng, Zhangquan Chen, Li Yu, Kai Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
相关 Paper
- Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-ResolutionKaram Park, Jae Woong Soh, Nam Ik ChoAAAI 2025 · 被引用 20 次
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran 等ACM MM 2025 · 被引用 3 次
- Recursive Generalization Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong 等ICLR 2024 · 被引用 81 次
- Scaling Attention via Feature SparsityYan Xie, Tiansheng Wen, Tangda Huang, Bo Chen 等ICLR 2026 · 被引用 3 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
