Neural Network Architecture Beyond Width and Depth
Shijun Zhang, Zuowei Shen, Haizhao Yang
Abstract
This paper proposes a new neural network architecture by introducing an additional dimension called height beyond width and depth. Neural network architectures with height, width, and depth as hyper-parameters are called three-dimensional architectures. It is shown that neural networks with three-dimensional architectures are significantly more expressive than the ones with two-dimensional architectures (those with only width and depth as hyper-parameters), e.g., standard fully connected networks. The new network architecture is constructed recursively via a nested structure, and hence we call a network with the new architecture nested network (NestNet). A NestNet of height is built with each hidden neuron activated by a NestNet of height . When , a NestNet degenerates to a standard network with a two-dimensional architecture. It is proved by construction that height- ReLU NestNets with parameters can approximate -Lipschitz continuous functions on with an error , while the optimal approximation error of standard ReLU networks with parameters is . Furthermore, such a result is extended to generic continuous functions on with the approximation error characterized by the modulus of continuity. Finally, we use numerical experimentation to show the advantages of the super-approximation power of ReLU NestNets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- On Enhancing Expressive Power via Compositions of Single Fixed-Size ReLU NetworkShijun Zhang, Jianfeng Lu, Hongkai ZhaoICML 2023 · 9 citations
- Characterizing ResNet's Universal Approximation CapabilityChenghao Liu, Enming Liang, Minghua ChenICML 2024
- KAN: Kolmogorov-Arnold NetworksZiming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle et al.ICLR 2025
Builds on6
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 156 citations
- Task Adaptive Parameter Sharing for Multi-Task LearningMatthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran et al.CVPR 2022 · 61 citations
- Revisiting Parameter Sharing for Automatic Neural Channel Number SearchJiaxing Wang, Haoli Bai, Jiaxiang Wu, Xupeng Shi et al.NeurIPS 2020 · 29 citations
- Deep Network Approximation in Terms of Intrinsic ParametersZuowei Shen, Haizhao Yang, Shijun ZhangICML 2022 · 13 citations
Related papers
- Better Neural Network Expressivity: Subdividing the SimplexEgor Bakaev, Florestan Brunck, Christoph Hertrich, Jack Stade et al.STOC 2026 · 16 citations
- Sharp Representation Theorems for ReLU Networks with Precise Dependence on DepthGuy Bresler, Dheeraj NagarajNeurIPS 2020 · 27 citations
- ReLU Network with Width d+O(1) Can Achieve Optimal Approximation RateChenghao Liu, Minghua ChenICML 2024 · 3 citations
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- On the Number of Linear Regions of Convolutional Neural NetworksHuan Xiong, Lei Huang, Mengyang Yu, Li Liu et al.ICML 2020 · 80 citations
