PolaFormer: Polarity-aware Linear Attention for Vision Transformers
Weikang Meng, Yadan Luo, Xin Li, Dongmei Jiang, Zheng Zhang
摘要
Linear attention has emerged as a promising alternative to softmax-based attention, leveraging kernelized feature maps to reduce complexity from quadratic to linear in sequence length. However, the non-negative constraint on feature maps and the relaxed exponential function used in approximation lead to significant information loss compared to the original query-key dot products, resulting in less discriminative attention maps with higher entropy. To address the missing interactions driven by negative values in query-key pairs, we propose a polarity-aware linear attention mechanism that explicitly models both same-signed and opposite-signed query-key interactions, ensuring comprehensive coverage of relational information. Furthermore, to restore the spiky properties of attention maps, we provide a theoretical analysis proving the existence of a class of element-wise functions (with positive first and second derivatives) that can reduce entropy in the attention distribution. For simplicity, and recognizing the distinct contributions of each dimension, we employ a learnable power function for rescaling, allowing strong and weak attention signals to be effectively separated. Extensive experiments demonstrate that the proposed PolaFormer improves performance on various vision tasks, enhancing both expressiveness and efficiency by up to 4.6%. Code is available at https://github.com/ZacharyMeng/PolaFormer . Figure 1: Attention weight visualization. Unlike prior linear attention approaches ((Katharopoulos et al., 2020) the 3rd and (Han et al., 2023a) 4th plots) that generate uniform responses, the proposed PolaFormer captures a more accurate query-key interaction with lower entropy, closely resembling softmax while maintaining linear complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LaplacianFormer: Rethinking Linear Attention with Laplacian KernelZhe Feng, Sen Lian, Changwei Wang, Muyang Zhang 等ICLR 2026 · 被引用 3 次
- Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language PretrainingWeijun Zhuang, Yuqing Huang, Weikang Meng, Xin Li 等CVPR 2026 · 被引用 3 次
- HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution EditingYuyao Zhang, Alexander Huang-Menders, Yu-Wing TaiCVPR 2026 · 被引用 2 次
- NormDirection: Restoring the Missing Query Norm in Vision Linear AttentionWeikang Meng, Yadan Luo, Liangyu Huo, Yingjian Li 等ICML 2026 · 被引用 2 次
- Vision Transformers Are Circulant Attention LearnersDongchen Han, Tianyu Li, Ziyi Wang, Gao HuangAAAI 2026 · 被引用 2 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
相关 Paper
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li 等ICLR 2022 · 被引用 303 次
- Rectifying Magnitude Neglect in Linear AttentionQihang Fan, Huaibo Huang, Yuang Ai, Ran HeICCV 2025 · 被引用 14 次
- PolySketchFormer: Fast Transformers via Sketching Polynomial KernelsPraneeth Kacham, Vahab Mirrokni, Peilin ZhongICML 2024 · 被引用 27 次
- Linear Complexity Randomized Self-attention MechanismLin Zheng, Chong Wang, Lingpeng KongICML 2022 · 被引用 39 次
- Flowformer: Linearizing Transformers with Conservation FlowsHaixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang 等ICML 2022 · 被引用 130 次
