Approximation Error Upper and Lower Bounds for Hölder Class with Transformers
Xin He, Yuling Jiao, Xiliang Lu, Jerry Yang
摘要
We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and residual connections. We prove that a Transformer network composed of at most blocks can approximate any bounded Hölder function with -dimensional input and smoothness under any accuracy . In the case of approximation lower bounds, leveraging the VC-dimension upper bound, we are the first to rigorously prove that Transformers demand for at least blocks to achieve the approximation accuracy. As a final step, we extend the derived results for standard Transformers to a general regression task and establish the corresponding excess risk rates demonstrating Transformers' empirical effectiveness in real-world settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 被引用 156 次
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 被引用 154 次
相关 Paper
- Approximation Bounds for Transformer Networks with Application to RegressionYuling Jiao, Yanming Lai, Defeng Sun, Yang Wang 等ICML 2026 · 被引用 6 次
- The Effect of Attention Head Count on Transformer ApproximationPenghao Yu, Haotian Jiang, Zeyu Bao, Ruoxi Yu 等ICLR 2026 · 被引用 5 次
- Efficient and Minimax Optimal In-context Nonparametric Regression with TransformersMichelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma 等ICML 2026 · 被引用 6 次
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 被引用 31 次
- In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separationFrank Cole, Yuxuan Zhao, Yulong Lu, Tianhao ZhangNeurIPS 2025 · 被引用 1 次
