Approximation Bounds for Transformer Networks with Application to Regression
Yuling Jiao, Yanming Lai, Defeng Sun, Yang Wang, Bokai Yan
Abstract
We explore the approximation capabilities of Transformer networks for Hölder and Sobolev functions, and apply these results to address nonparametric regression estimation with dependent observations. First, we establish novel upper bounds for standard Transformer networks approximating sequence-to-sequence mappings whose component functions are Hölder continuous with smoothness index γ P p0, 1s. To achieve an approximation error ε under the L p -norm for p P r1, 8s, it suffices to use a fixed-depth Transformer network whose total number of parameters scales as ε ´dxnγ . This result not only extends existing findings to include the case p " 8, but also matches the best known upper bounds on number of parameters previously obtained for fixed-depth FNNs and RNNs. Similar bounds are also derived for Sobolev functions. Second, we derive explicit convergence rates for the nonparametric regression problem under various β-mixing data assumptions, which allow the dependence between observations to weaken over time. Our bounds on the sample complexity impose no constraints on weight magnitudes. Lastly, we propose a novel proof strategy to establish approximation bounds, inspired by the Kolmogorov-Arnold representation theorem. We show that if the self-attention layer in a Transformer can perform column averaging, the network can approximate sequence-to-sequence Hölder functions, offering new insights into the interpretability of self-attention mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Prompt Tuning Transformers for Data MemorizationHaiyu Wang, Yuanyuan LinNeurIPS 2025 · 4 citations
- Posterior Contraction for Sparse Neural Networks in Besov Spaces with Intrinsic DimensionalityKyeongwon Lee, Lizhen Lin, Jaewoo Park, Seonghyun JeongNeurIPS 2025 · 4 citations
- Approximation Error Upper and Lower Bounds for Hölder Class with TransformersXin He, Yuling Jiao, Xiliang Lu, Jerry YangICML 2026
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou et al.NeurIPS 2023 · 182 citations
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat et al.NeurIPS 2020 · 111 citations
Related papers
- Efficient and Minimax Optimal In-context Nonparametric Regression with TransformersMichelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma et al.ICML 2026 · 6 citations
- Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional InputShokichi Takakura, Taiji SuzukiICML 2023 · 32 citations
- The Effect of Attention Head Count on Transformer ApproximationPenghao Yu, Haotian Jiang, Zeyu Bao, Ruoxi Yu et al.ICLR 2026 · 5 citations
- Approximation Rate of the Transformer Architecture for Sequence ModelingHaotian Jiang, Qianxiao LiNeurIPS 2024 · 32 citations
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 31 citations
