Elliptical Attention
Stefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan Nguyen
Abstract
Pairwise dot-product self-attention is key to the success of transformers that achieve state-of-the-art performance across a variety of applications in language and vision. This dot-product self-attention computes attention weights among the input tokens using Euclidean distance, which makes the model prone to representation collapse and vulnerable to contaminated samples. In this paper, we propose using a Mahalanobis distance metric for computing the attention weights to stretch the underlying feature space in directions of high contextual relevance. In particular, we define a hyper-ellipsoidal neighborhood around each query to increase the attention weights of the tokens lying in the contextually important directions. We term this novel class of attention Elliptical Attention. Our Elliptical Attention provides two benefits: 1) reducing representation collapse and 2) enhancing the model's robustness as Elliptical Attention pays more attention to contextually relevant information rather than focusing on some small subset of informative features. We empirically demonstrate the advantages of Elliptical Attention over the baseline dot-product attention and state-of-the-art attention methods on various practical tasks, including object classification, image segmentation, and language modeling across different data modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b2fed5c-b34e-4243-a1b9-c631fff54e59Cited by top-tier papers7
- Tropical Attention: Neural Algorithmic Reasoning for Combinatorial AlgorithmsBaran Hashemi, Kurt Pasque, Christopher Teska, Ruriko YoshidaNeurIPS 2025 · 14 citations
- Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component AnalysisRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 11 citations
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 1 citation
- Transformer Meets Twicing: Harnessing Unattended Residual InformationLaziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Equivariant Neural Functional Networks for TransformersHoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran Huu et al.ICLR 2025
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Coneheads: Hierarchy Aware AttentionAlbert Tseng, Tao Yu, Toni J. B. Liu, Christopher De SaNeurIPS 2023 · 8 citations
- Geodesic Self-Attention for 3D Point CloudsZhengyu Li, Xuan Tang, Zihao Xu, Xihao Wang et al.NeurIPS 2022 · 18 citations
- Rectifying Magnitude Neglect in Linear AttentionQihang Fan, Huaibo Huang, Yuang Ai, Ran HeICCV 2025 · 14 citations
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan et al.ICML 2021 · 399 citations
- Neighborhood Attention TransformerAli Hassani, Steven Walton, Jiachen Li, Shen Li et al.CVPR 2023
