Approximation Rate of the Transformer Architecture for Sequence Modeling
Haotian Jiang, Qianxiao Li
Abstract
The Transformer architecture is widely applied in sequence modeling applications, yet the theoretical understanding of its working principles remains limited. In this work, we investigate the approximation rate for single-layer Transformers with one head. We consider a class of non-linear relationships and identify a novel notion of complexity measures to establish an explicit Jackson-type approximation rate estimate for the Transformer. This rate reveals the structural properties of the Transformer and suggests the types of sequential relationships it is best suited for approximating. In particular, the results on approximation rates enable us to concretely analyze the differences between the Transformer and classical sequence modeling methods, such as recurrent neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c5377c9-5baf-4cf3-a55d-9b4eda1398d5Cited by top-tier papers14
- High-Order Flow Matching: Unified Framework and Sharp Statistical RatesMaojiang Su, Jerry Yao-Chieh Hu, Yi-Chen Lee, Ning Zhu et al.NeurIPS 2025 · 9 citations
- Universal Approximation with Softmax AttentionJerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu et al.ICML 2026 · 9 citations
- Approximation Bounds for Transformer Networks with Application to RegressionYuling Jiao, Yanming Lai, Defeng Sun, Yang Wang et al.ICML 2026 · 6 citations
- The Effect of Attention Head Count on Transformer ApproximationPenghao Yu, Haotian Jiang, Zeyu Bao, Ruoxi Yu et al.ICLR 2026 · 5 citations
- Prompt Tuning Transformers for Data MemorizationHaiyu Wang, Yuanyuan LinNeurIPS 2025 · 4 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
Related papers
- Understanding the Expressive Power and Mechanisms of Transformer for Sequence ModelingMingze Wang, Weinan ENeurIPS 2024 · 32 citations
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 8 citations
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical PerspectiveAlireza Mousavi-Hosseini, Clayton Sanford, Denny Wu, Murat A. ErdogduNeurIPS 2025 · 6 citations
- In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separationFrank Cole, Yuxuan Zhao, Yulong Lu, Tianhao ZhangNeurIPS 2025 · 1 citation
- Transformers, parallel computation, and logarithmic depthClayton Sanford, Daniel Hsu, Matus TelgarskyICML 2024 · 64 citations
