Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural Networks
Julia B. Nakhleh, Robert D. Nowak
摘要
Overparameterized neural networks can interpolate a given dataset in many different ways, prompting the fundamental question: which among these solutions should we prefer, and what explicit regularization strategies will provably yield these solutions? This paper addresses the challenge of finding the sparsest interpolating ReLU network--i.e., the network with the fewest nonzero parameters or neurons--a goal with wide-ranging implications for efficiency, generalization, interpretability, theory, and model compression. Unlike post hoc pruning approaches, we propose a continuous, almost-everywhere differentiable training objective whose global minima are guaranteed to correspond to the sparsest single-hidden-layer ReLU networks that fit the data. This result marks a conceptual advance: it recasts the combinatorial problem of sparse interpolation as a smooth optimization task, potentially enabling the use of gradient-based training methods. Our objective is based on minimizing quasinorms of the weights for , a classical sparsity-promoting strategy in finite-dimensional settings. However, applying these ideas to neural networks presents new challenges: the function class is infinite-dimensional, and the weights are learned using a highly nonconvex objective. We prove that, under our formulation, global minimizers correspond exactly to sparsest solutions. Our work lays a foundation for understanding when and how continuous sparsity-inducing objectives can be leveraged to recover sparse networks through training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 被引用 152 次
- Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer NetworksMert Pilanci, Tolga ErgenICML 2020 · 被引用 142 次
- Network size and size of the weights in memorization with two-layers neural networksSébastien Bubeck, Ronen Eldan, Yin Tat Lee, Dan MikulincerNeurIPS 2020 · 被引用 28 次
- Penalising the biases in norm regularisation enforces sparsityEtienne Boursier, Nicolas FlammarionNeurIPS 2023 · 被引用 21 次
- A New Neural Kernel Regime: The Inductive Bias of Multi-Task LearningJulia B. Nakhleh, Joseph Shenouda, Robert D. NowakNeurIPS 2024 · 被引用 3 次
相关 Paper
- Optimal Sets and Solution Paths of ReLU NetworksAaron Mishkin, Mert PilanciICML 2023 · 被引用 7 次
- Does a sparse ReLU network training problem always admit an optimum ?Quoc-Tung Le, Rémi Gribonval, Elisa RicciettiNeurIPS 2023 · 被引用 5 次
- Minimum norm interpolation by perceptra: Explicit regularization and implicit biasJiyoung Park, Ian Pelakh, Stephan WojtowytschNeurIPS 2023 · 被引用 5 次
- Global Optimality Beyond Two Layers: Training Deep ReLU Networks via Convex ProgramsTolga Ergen, Mert PilanciICML 2021 · 被引用 35 次
- spred: Solving L1 Penalty with SGDLiu Ziyin, Zihao WangICML 2023 · 被引用 23 次
