Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
Rachel S. Y. Teo, Tan M. Nguyen
摘要
The remarkable success of transformers in sequence modeling tasks, spanning various applications in natural language processing and computer vision, is attributed to the critical role of self-attention. Similar to the development of most deep learning models, the construction of these attention mechanisms relies on heuristics and experience. In our work, we derive self-attention from kernel principal component analysis (kernel PCA) and show that self-attention projects its query vectors onto the principal component axes of its key matrix in a feature space. We then formulate the exact formula for the value matrix in self-attention, theoretically and empirically demonstrating that this value matrix captures the eigenvectors of the Gram matrix of the key vectors in self-attention. Leveraging our kernel PCA framework, we propose Attention with Robust Principal Components (RPC-Attention), a novel class of robust attention that is resilient to data contamination. We empirically demonstrate the advantages of RPC-Attention over softmax attention on the ImageNet-1K object classification, WikiText-103 language modeling, and ADE20K image segmentation task. The code is publicly available at https://github.com/rachtsy/KPCA_code .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Elliptical AttentionStefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan NguyenNeurIPS 2024 · 被引用 12 次
- Transformer Meets Twicing: Harnessing Unattended Residual InformationLaziz U. Abdullaev, Tan Minh NguyenICLR 2025
- Rethinking PCA Through DualityJan Quan, Johan A. K. Suykens, Panagiotis PatrinosNeurIPS 2025
- Equivariant Neural Functional Networks for TransformersHoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran Huu 等ICLR 2025
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
它引用的顶会 Paper43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
相关 Paper
- Designing Robust Transformers using Robust Kernel Density EstimationXing Han, Tongzheng Ren, Tan Nguyen, Khai Nguyen 等NeurIPS 2023 · 被引用 14 次
- Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression PerspectiveQishuai Wen, Chun-Guang LiNeurIPS 2024 · 被引用 8 次
- Improving Robustness of Vision Transformers by Reducing Sensitivity to Patch CorruptionsYong Guo, David Stutz, Bernt SchieleCVPR 2023
- Can CNNs Be More Robust Than Transformers?Zeyu Wang, Yutong Bai, Yuyin Zhou, Cihang XieICLR 2023 · 被引用 14 次
- A Primal-Dual Framework for Transformers and Neural NetworksTan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi 等ICLR 2023 · 被引用 2 次
