Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think
Christian H. X. Ali Mehmeti-Göpel, Jan Disselhoff
摘要
We perform an empirical study of the behaviour of deep networks when fully linearizing some of its feature channels through a sparsity prior on the overall number of nonlinear units in the network. In experiments on image classification and machine translation tasks, we investigate how much we can simplify the network function towards linearity before performance collapses. First, we observe a significant performance gap when reducing nonlinearity in the network function early on as opposed to late in training, in-line with recent observations on the time-evolution of the data-dependent NTK. Second, we find that after training, we are able to linearize a significant number of nonlinear units while maintaining a high performance, indicating that much of a network's expressivity remains unused but helps gradient descent in early stages of training. To characterize the depth of the resulting partially linearized network, we introduce a measure called average path length, representing the average number of active nonlinearities encountered along a path in the network graph. Under sparsity pressure, we find that the remaining nonlinear units organize into distinct structures, forming core-networks of near constant effective depth and width, which in turn depend on task difficulty.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Till the Layers Collapse: Compressing a Deep Neural Network Through the Lenses of Batch Normalization LayersZhu Liao, Nour Hezbri, Victor Quétu, Van-Tam Nguyen 等AAAI 2025 · 被引用 3 次
- LaCoOT: Layer Collapse through Optimal TransportVictor Quétu, Zhu Liao, Nour Hezbri, Fabio Pizzati 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper7
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Drawing Early-Bird Tickets: Toward More Efficient Training of Deep NetworksHaoran You, Chaojian Li, Pengfei Xu, Yonggan Fu 等ICLR 2020 · 被引用 282 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Sanity-Checking Pruning Methods: Random Tickets can Win the JackpotJingtong Su, Yihang Chen, Tianle Cai, Tianhao Wu 等NeurIPS 2020 · 被引用 100 次
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 被引用 26 次
相关 Paper
- Frivolous Units: Wider Networks Are Not Really That WideStephen Casper, Xavier Boix, Vanessa D'Amario, Ling Guo 等AAAI 2021 · 被引用 20 次
- Finding Lottery Tickets in Vision Models via Data-Driven Spectral Foresight PruningLeonardo Iurada, Marco Ciccone, Tatiana TommasiCVPR 2024
- Phase Collapse in Neural NetworksFlorentin Guth, John Zarka, Stéphane MallatICLR 2022 · 被引用 8 次
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
- Learning sparse features can lead to overfitting in neural networksLeonardo Petrini, Francesco Cagnetta, Eric Vanden-Eijnden, Matthieu WyartNeurIPS 2022 · 被引用 47 次
