Sharp Representation Theorems for ReLU Networks with Precise Dependence on Depth
Guy Bresler, Dheeraj Nagaraj
2020Year
27Citations
4Top-tier citations
Abstract
We prove sharp dimension-free representation results for neural networks with ReLU layers under square loss for a class of functions defined in the paper. These results capture the precise benefits of depth in the following sense:
- The rates for representing the class of functions via ReLU layers is sharp up to constants, as shown by matching lower bounds.
- For each , and as grows the class of functions contains progressively less smooth functions.
- If , then the approximation rate for the class achieved by depth networks is strictly worse than that achieved by depth networks. This constitutes a fine-grained characterization of the representation power of feedforward networks of arbitrary depth and number of neurons , in contrast to existing representation results which either require growing quickly with or assume that the function being represented is highly smooth. In the latter case similar rates can be obtained with a single nonlinear layer. Our results confirm the prevailing hypothesis that deeper networks are better at representing less smooth functions, and indeed, the main technical novelty is to fully exploit the fact that deep networks can produce highly oscillatory functions with few activation functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bef250f-126b-431d-b4d9-20946b3cf138Cited by top-tier papers4
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler et al.NeurIPS 2021 · 74 citations
- On the Optimal Memorization Power of ReLU Neural NetworksGal Vardi, Gilad Yehudai, Ohad ShamirICLR 2022 · 42 citations
- On the Representation of Solutions to Elliptic PDEs in Barron SpacesZiang Chen, Jianfeng Lu, Yulong LuNeurIPS 2021 · 42 citations
- Theoretical limitations of multi-layer TransformerLijie Chen, Binghui Peng, Hongxun WuFOCS 2025 · 2 citations
Builds on2
- Depth-Width Trade-offs for ReLU Networks via Sharkovsky's TheoremVaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis Panageas, Xiao WangICLR 2020 · 24 citations
- Better depth-width trade-offs for neural networks through the lens of dynamical systemsVaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis PanageasICML 2020 · 15 citations
Related papers
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 130 citations
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 156 citations
- On Enhancing Expressive Power via Compositions of Single Fixed-Size ReLU NetworkShijun Zhang, Jianfeng Lu, Hongkai ZhaoICML 2023 · 9 citations
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao et al.NeurIPS 2020 · 61 citations
