Spectral Conditioning of Attention Improves Transformer Performance
Hemanth Saratchandran, Simon Lucey
Abstract
We present a theoretical analysis of the Jacobian of an attention block within a transformer, showing that it is governed by the query, key, and value projections that define the attention mechanism. Leveraging this insight, we introduce a method that systematically alters the spectral properties of each attention layer to reduce the Jacobian's condition number, thereby improving the overall conditioning of the attention layers within a transformer network. We empirically show that this improved Jacobian conditioning translates to enhanced performance in practice. Our approach is simple, broadly applicable, and can be easily integrated as a drop-in replacement for a wide range of existing attention mechanisms. We validate its effectiveness across diverse transformer architectures and tasks, demonstrating consistent improvements in performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b47c568-ed28-478f-b929-df2201cb94ccCited by top-tier papers2
- SineProject: Machine Unlearning for Stable Vision-Language AlignmentArpit Garg, Hemanth Saratchandran, Simon LuceyCVPR 2026 · 2 citations
- CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News GenerationZhao Tong, Chunlin Gong, Yiping Zhang, Haichao Shi et al.ICML 2026
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski et al.NeurIPS 2021 · 692 citations
Related papers
- Conditioned Initialization for AttentionHemanth Saratchandran, Simon LuceyICLR 2026
- Enhancing Transformers Through Conditioned Embedded TokensHemanth Saratchandran, Simon LuceyICCV 2025
- Jump Self-attention: Capturing High-order Statistics in TransformersHaoyi Zhou, Siyang Xiao, Shanghang Zhang, Jieqi Peng et al.NeurIPS 2022 · 4 citations
- Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention LayersThiziri Nait Saada, Alireza Naderi, Jared TannerICML 2025
- Self-Adjust SoftmaxChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi et al.EMNLP 2025
