Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During Training
Max Milkert, David Hyde, Forrest J. Laine
Abstract
In a neural network with ReLU activations, the number of piecewise linear regions in the output can grow exponentially with depth. However, this is highly unlikely to happen when the initial parameters are sampled randomly, which therefore often leads to the use of networks that are unnecessarily large. To address this problem, we introduce a novel parameterization of the network that restricts its weights so that a depth d network produces exactly 2 d linear regions at initialization and maintains those regions throughout training under the parameterization. This approach allows us to learn approximations of convex, onedimensional functions that are several orders of magnitude more accurate than their randomly initialized counterparts. We further demonstrate a preliminary extension of our construction to multidimensional and non-convex functions, allowing the technique to replace traditional dense layers in various architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dc5c451-c224-48ec-a850-5bbcabbbec02Builds on3
- Reverse-engineering deep ReLU networksDavid Rolnick, Konrad P. KordingICML 2020 · 121 citations
- Unsupervised Representation Learning via Neural Activation CodingYookoon Park, Sangho Lee, Gunhee Kim, David M. BleiICML 2021 · 8 citations
- Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU NetworksWenlin Chen, Hong GeNeurIPS 2024 · 5 citations
Related papers
- On the Expected Complexity of Maxout NetworksHanna Tseran, Guido MontúfarNeurIPS 2021 · 19 citations
- Most Activation Functions Can Win the Lottery Without Excessive DepthRebekka BurkholzNeurIPS 2022 · 27 citations
- Deep Network Approximation in Terms of Intrinsic ParametersZuowei Shen, Haizhao Yang, Shijun ZhangICML 2022 · 13 citations
- Sharp Representation Theorems for ReLU Networks with Precise Dependence on DepthGuy Bresler, Dheeraj NagarajNeurIPS 2020 · 27 citations
- On the Number of Linear Regions of Convolutional Neural NetworksHuan Xiong, Lei Huang, Mengyang Yu, Li Liu et al.ICML 2020 · 80 citations
