Universal Approximation with Softmax Attention
Jerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu, Han Liu
Abstract
We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on compact domains. Our main technique is a new interpolation-based method for analyzing attention’s internal mechanism. This leads to our key insight: self-attention is able to approximate a generalized version of ReLU to arbitrary precision, and hence subsumes many known universal approximators. Building on these, we show that two-layer multi-head attention or even one-layer multi-head attention followed by a softmax function suffices as a sequence-to-sequence universal approximator. In contrast, prior works rely on feed-forward networks to establish universal approximation in Transformers. Furthermore, we extend our techniques to show that, (softmax-)attention-only layers are capable of approximating gradient descent in-context. We believe these techniques hold independent interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 999895f6-f725-43f5-af94-cfd4fa30dc6eCited by top-tier papers11
- Attention Mechanism, Max-Affine Partition, and Universal ApproximationHude Liu, Jerry Yao-Chieh Hu, Zhao Song, Han LiuNeurIPS 2025 · 12 citations
- High-Order Flow Matching: Unified Framework and Sharp Statistical RatesMaojiang Su, Jerry Yao-Chieh Hu, Yi-Chen Lee, Ning Zhu et al.NeurIPS 2025 · 9 citations
- In-Context Algorithm Emulation in Fixed-Weight TransformersJerry Yao-Chieh Hu, Hude Liu, Jennifer Yuntong Zhang, Han LiuICLR 2026 · 7 citations
- Prompt Tuning Transformers for Data MemorizationHaiyu Wang, Yuanyuan LinNeurIPS 2025 · 4 citations
- A unified framework for establishing the universal approximation of transformer-type architecturesJingpu Cheng, Ting Lin, Zuowei Shen, Qianxiao LiNeurIPS 2025 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- Minimum Width for Universal ApproximationSejun Park, Chulhee Yun, Jaeho Lee, Jinwoo ShinICLR 2021 · 148 citations
- Low-Rank Bottleneck in Multi-head Attention ModelsSrinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi et al.ICML 2020 · 130 citations
Related papers
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 31 citations
- Prompting a Pretrained Transformer Can Be a Universal ApproximatorAleksandar Petrov, Philip Torr, Adel BibiICML 2024 · 19 citations
- In-Context Universal Approximation, Compositional Generalization, and Algorithm EmulationJerry Yao-Chieh Hu, Hong-Yu Chen, Po-Chiao Lin, Maojiang Su et al.ICML 2026
- Theory, Analysis, and Best Practices for Sigmoid Self-AttentionJason Ramapuram, Federico Danieli, Eeshan Gunesh Dhekane, Floris Weers et al.ICLR 2025
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat et al.NeurIPS 2020 · 111 citations
