Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
Rachel S. Y. Teo, Tan M. Nguyen
Abstract
The remarkable success of transformers in sequence modeling tasks, spanning various applications in natural language processing and computer vision, is attributed to the critical role of self-attention. Similar to the development of most deep learning models, the construction of these attention mechanisms relies on heuristics and experience. In our work, we derive self-attention from kernel principal component analysis (kernel PCA) and show that self-attention projects its query vectors onto the principal component axes of its key matrix in a feature space. We then formulate the exact formula for the value matrix in self-attention, theoretically and empirically demonstrating that this value matrix captures the eigenvectors of the Gram matrix of the key vectors in self-attention. Leveraging our kernel PCA framework, we propose Attention with Robust Principal Components (RPC-Attention), a novel class of robust attention that is resilient to data contamination. We empirically demonstrate the advantages of RPC-Attention over softmax attention on the ImageNet-1K object classification, WikiText-103 language modeling, and ADE20K image segmentation task. The code is publicly available at https://github.com/rachtsy/KPCA_code .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0f8bb80-03a2-4cd8-b2f9-bfe12b59f7a0Cited by top-tier papers8
- Elliptical AttentionStefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan NguyenNeurIPS 2024 · 12 citations
- Transformer Meets Twicing: Harnessing Unattended Residual InformationLaziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Rethinking PCA Through DualityJan Quan, Johan A. K. Suykens, Panagiotis PatrinosNeurIPS 2025
- Equivariant Neural Functional Networks for TransformersHoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran Huu et al.ICLR 2025
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
Builds on43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
Related papers
- Designing Robust Transformers using Robust Kernel Density EstimationXing Han, Tongzheng Ren, Tan Nguyen, Khai Nguyen et al.NeurIPS 2023 · 14 citations
- Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression PerspectiveQishuai Wen, Chun-Guang LiNeurIPS 2024 · 8 citations
- Improving Robustness of Vision Transformers by Reducing Sensitivity to Patch CorruptionsYong Guo, David Stutz, Bernt SchieleCVPR 2023
- Can CNNs Be More Robust Than Transformers?Zeyu Wang, Yutong Bai, Yuyin Zhou, Cihang XieICLR 2023 · 14 citations
- A Primal-Dual Framework for Transformers and Neural NetworksTan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi et al.ICLR 2023 · 2 citations
