Neural Network Architecture Beyond Width and Depth
Shijun Zhang, Zuowei Shen, Haizhao Yang
摘要
This paper proposes a new neural network architecture by introducing an additional dimension called height beyond width and depth. Neural network architectures with height, width, and depth as hyper-parameters are called three-dimensional architectures. It is shown that neural networks with three-dimensional architectures are significantly more expressive than the ones with two-dimensional architectures (those with only width and depth as hyper-parameters), e.g., standard fully connected networks. The new network architecture is constructed recursively via a nested structure, and hence we call a network with the new architecture nested network (NestNet). A NestNet of height is built with each hidden neuron activated by a NestNet of height . When , a NestNet degenerates to a standard network with a two-dimensional architecture. It is proved by construction that height- ReLU NestNets with parameters can approximate -Lipschitz continuous functions on with an error , while the optimal approximation error of standard ReLU networks with parameters is . Furthermore, such a result is extended to generic continuous functions on with the approximation error characterized by the modulus of continuity. Finally, we use numerical experimentation to show the advantages of the super-approximation power of ReLU NestNets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On Enhancing Expressive Power via Compositions of Single Fixed-Size ReLU NetworkShijun Zhang, Jianfeng Lu, Hongkai ZhaoICML 2023 · 被引用 9 次
- Characterizing ResNet's Universal Approximation CapabilityChenghao Liu, Enming Liang, Minghua ChenICML 2024
- KAN: Kolmogorov-Arnold NetworksZiming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle 等ICLR 2025
它引用的顶会 Paper6
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 被引用 156 次
- Task Adaptive Parameter Sharing for Multi-Task LearningMatthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran 等CVPR 2022 · 被引用 61 次
- Revisiting Parameter Sharing for Automatic Neural Channel Number SearchJiaxing Wang, Haoli Bai, Jiaxiang Wu, Xupeng Shi 等NeurIPS 2020 · 被引用 29 次
- Deep Network Approximation in Terms of Intrinsic ParametersZuowei Shen, Haizhao Yang, Shijun ZhangICML 2022 · 被引用 13 次
相关 Paper
- Better Neural Network Expressivity: Subdividing the SimplexEgor Bakaev, Florestan Brunck, Christoph Hertrich, Jack Stade 等STOC 2026 · 被引用 16 次
- Sharp Representation Theorems for ReLU Networks with Precise Dependence on DepthGuy Bresler, Dheeraj NagarajNeurIPS 2020 · 被引用 27 次
- ReLU Network with Width d+O(1) Can Achieve Optimal Approximation RateChenghao Liu, Minghua ChenICML 2024 · 被引用 3 次
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- On the Number of Linear Regions of Convolutional Neural NetworksHuan Xiong, Lei Huang, Mengyang Yu, Li Liu 等ICML 2020 · 被引用 80 次
