Lune

ICML2026Top-tier venue

Approximation Error Upper and Lower Bounds for Hölder Class with Transformers

Xin He, Yuling Jiao, Xiliang Lu, Jerry Yang

2026Year

Abstract

We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and residual connections. We prove that a Transformer network composed of at most O(ε−d0/α)\mathcal{O}(\varepsilon^{-{d_{0}}/{\alpha}}) blocks can approximate any bounded Hölder function with d0d_{0}-dimensional input and smoothness α∈(0,1]\alpha\in(0,1] under any accuracy ε>0\varepsilon>0. In the case of approximation lower bounds, leveraging the VC-dimension upper bound, we are the first to rigorously prove that Transformers demand for at least Ω(ε−d0/(4α))\Omega(\varepsilon^{-{d_{0}}/({4\alpha})}) blocks to achieve the ε\varepsilon approximation accuracy. As a final step, we extend the derived results for standard Transformers to a general regression task and establish the corresponding excess risk rates demonstrating Transformers' empirical effectiveness in real-world settings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3dd47c04-d519-41ee-9f94-cd7f2e1e3b03

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines