Sharp Representation Theorems for ReLU Networks with Precise Dependence on Depth
Guy Bresler, Dheeraj Nagaraj
2020年份
27被引次数
4顶会引用
摘要
We prove sharp dimension-free representation results for neural networks with ReLU layers under square loss for a class of functions defined in the paper. These results capture the precise benefits of depth in the following sense:
- The rates for representing the class of functions via ReLU layers is sharp up to constants, as shown by matching lower bounds.
- For each , and as grows the class of functions contains progressively less smooth functions.
- If , then the approximation rate for the class achieved by depth networks is strictly worse than that achieved by depth networks. This constitutes a fine-grained characterization of the representation power of feedforward networks of arbitrary depth and number of neurons , in contrast to existing representation results which either require growing quickly with or assume that the function being represented is highly smooth. In the latter case similar rates can be obtained with a single nonlinear layer. Our results confirm the prevailing hypothesis that deeper networks are better at representing less smooth functions, and indeed, the main technical novelty is to fully exploit the fact that deep networks can produce highly oscillatory functions with few activation functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler 等NeurIPS 2021 · 被引用 74 次
- On the Optimal Memorization Power of ReLU Neural NetworksGal Vardi, Gilad Yehudai, Ohad ShamirICLR 2022 · 被引用 42 次
- On the Representation of Solutions to Elliptic PDEs in Barron SpacesZiang Chen, Jianfeng Lu, Yulong LuNeurIPS 2021 · 被引用 42 次
- Theoretical limitations of multi-layer TransformerLijie Chen, Binghui Peng, Hongxun WuFOCS 2025 · 被引用 2 次
它引用的顶会 Paper2
- Depth-Width Trade-offs for ReLU Networks via Sharkovsky's TheoremVaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis Panageas, Xiao WangICLR 2020 · 被引用 24 次
- Better depth-width trade-offs for neural networks through the lens of dynamical systemsVaggos Chatziafratis, Sai Ganesh Nagarajan, Ioannis PanageasICML 2020 · 被引用 15 次
相关 Paper
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 被引用 130 次
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 被引用 156 次
- On Enhancing Expressive Power via Compositions of Single Fixed-Size ReLU NetworkShijun Zhang, Jianfeng Lu, Hongkai ZhaoICML 2023 · 被引用 9 次
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao 等NeurIPS 2020 · 被引用 61 次
