A closer look at the approximation capabilities of neural networks
Kai Fong Ernest Chong
Abstract
The universal approximation theorem, in one of its most general versions, says that if we consider only continuous activation functions , then a standard feedforward neural network with one hidden layer is able to approximate any continuous multivariate function to any given approximation threshold , if and only if is non-polynomial. In this paper, we give a direct algebraic proof of the theorem. Furthermore we shall explicitly quantify the number of hidden units required for approximation. Specifically, if is compact, then a neural network with input units, output units, and a single hidden layer with hidden units (independent of and ), can uniformly approximate any polynomial function whose total degree is at most for each of its coordinate functions. In the general case that is any continuous function, we show there exists some (independent of ), such that hidden units would suffice to approximate . We also show that this uniform approximation property (UAP) still holds even under seemingly strong conditions imposed on the weights. We highlight several consequences: (i) For any , the UAP still holds if we restrict all non-bias weights in the last layer to satisfy (depending only on and ), such that the UAP still holds if we restrict all non-bias weights in the first layer to satisfy . (iii) If the non-bias weights in the first layer are fixed and randomly chosen from a suitable range, then the UAP holds with probability .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 901042c4-1250-493d-bbcf-41847f4722b6Cited by top-tier papers4
- Learning Hard Optimization Problems: A Data Generation PerspectiveJames Kotary, Ferdinando Fioretto, Pascal Van HentenryckNeurIPS 2021 · 45 citations
- Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal ApproximationLi'ang Li, Yifei Duan, Guanghua Ji, Yongqiang CaiICML 2023 · 20 citations
- Compiling to Linear NeuronsJoey Velez-Ginorio, Nada Amin, Konrad P. Kording, Steve ZdancewicPOPL 2026 · 1 citation
- Abstract Visual Reasoning: An Algebraic Approach for Solving Raven's Progressive MatricesJingyi Xu, Tushar Vaidya, Yufei Wu, Saket Chandra et al.CVPR 2023
Related papers
- Achieve the Minimum Width of Neural Networks for Universal ApproximationYongqiang CaiICLR 2023 · 4 citations
- Structure of universal formulasDmitry YarotskyNeurIPS 2023 · 2 citations
- Towards Lower Bounds on the Depth of ReLU Neural NetworksChristoph Hertrich, Amitabh Basu, Marco Di Summa, Martin SkutellaNeurIPS 2021 · 70 citations
- Universal approximation power of deep residual neural networks via nonlinear control theoryPaulo Tabuada, Bahman GharesifardICLR 2021 · 31 citations
- Optimal Minimum Width for the Universal Approximation of Continuously Differentiable Functions by Deep Narrow MLPsGeonho HwangNeurIPS 2025 · 2 citations
