Approximation Bounds for Transformer Networks with Application to Regression
Yuling Jiao, Yanming Lai, Defeng Sun, Yang Wang, Bokai Yan
摘要
We explore the approximation capabilities of Transformer networks for Hölder and Sobolev functions, and apply these results to address nonparametric regression estimation with dependent observations. First, we establish novel upper bounds for standard Transformer networks approximating sequence-to-sequence mappings whose component functions are Hölder continuous with smoothness index γ P p0, 1s. To achieve an approximation error ε under the L p -norm for p P r1, 8s, it suffices to use a fixed-depth Transformer network whose total number of parameters scales as ε ´dxnγ . This result not only extends existing findings to include the case p " 8, but also matches the best known upper bounds on number of parameters previously obtained for fixed-depth FNNs and RNNs. Similar bounds are also derived for Sobolev functions. Second, we derive explicit convergence rates for the nonparametric regression problem under various β-mixing data assumptions, which allow the dependence between observations to weaken over time. Our bounds on the sample complexity impose no constraints on weight magnitudes. Lastly, we propose a novel proof strategy to establish approximation bounds, inspired by the Kolmogorov-Arnold representation theorem. We show that if the self-attention layer in a Transformer can perform column averaging, the network can approximate sequence-to-sequence Hölder functions, offering new insights into the interpretability of self-attention mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Prompt Tuning Transformers for Data MemorizationHaiyu Wang, Yuanyuan LinNeurIPS 2025 · 被引用 4 次
- Posterior Contraction for Sparse Neural Networks in Besov Spaces with Intrinsic DimensionalityKyeongwon Lee, Lizhen Lin, Jaewoo Park, Seonghyun JeongNeurIPS 2025 · 被引用 4 次
- Approximation Error Upper and Lower Bounds for Hölder Class with TransformersXin He, Yuling Jiao, Xiliang Lu, Jerry YangICML 2026
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 被引用 154 次
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat 等NeurIPS 2020 · 被引用 111 次
相关 Paper
- Efficient and Minimax Optimal In-context Nonparametric Regression with TransformersMichelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma 等ICML 2026 · 被引用 6 次
- Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional InputShokichi Takakura, Taiji SuzukiICML 2023 · 被引用 32 次
- The Effect of Attention Head Count on Transformer ApproximationPenghao Yu, Haotian Jiang, Zeyu Bao, Ruoxi Yu 等ICLR 2026 · 被引用 5 次
- Approximation Rate of the Transformer Architecture for Sequence ModelingHaotian Jiang, Qianxiao LiNeurIPS 2024 · 被引用 32 次
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 被引用 31 次
