A Primal-Dual Framework for Transformers and Neural Networks
Tan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi, Richard G. Baraniuk, Stanley J. Osher
Abstract
Self-attention is key to the remarkable success of transformers in sequence modeling tasks including many applications in natural language processing and computer vision. Like neural network layers, these attention mechanisms are often developed by heuristics and experience. To provide a principled framework for constructing attention layers in transformers, we show that the self-attention corresponds to the support vector expansion derived from a support vector regression problem, whose primal formulation has the form of a neural network layer. Using our framework, we derive popular attention layers used in practice and propose two new attentions: 1) the Batch Normalized Attention (Attention-BN) derived from the batch normalization layer and 2) the Attention with Scaled Head (Attention-SH) derived from using less training data to fit the SVR model. We empirically demonstrate the advantages of the Attention-BN and Attention-SH in reducing head redundancy, increasing the model's accuracy, and improving the model's efficiency in a variety of practical applications including image and time-series classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7c0f389-5e35-486c-866d-73b9cdbd3e8aCited by top-tier papers10
- Max-Margin Token Selection in Attention MechanismDavoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet OymakNeurIPS 2023 · 67 citations
- Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensNeurIPS 2023 · 42 citations
- Elliptical AttentionStefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan NguyenNeurIPS 2024 · 12 citations
- Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component AnalysisRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 11 citations
- Towards Causal Foundation Model: on Duality between Optimal Balancing and AttentionJiaqi Zhang, Joel Jennings, Agrin Hilmkil, Nick Pawlowski et al.ICML 2024 · 9 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 18 citations
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 40 citations
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen et al.NeurIPS 2022 · 59 citations
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersLorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim SompolinskyNeurIPS 2024 · 12 citations
