Representation Costs of Linear Neural Networks: Analysis and Design
Zhen Dai, Mina Karzand, Nathan Srebro
摘要
For different parameterizations (mappings from parameters to predictors), we study the regularization cost in predictor space induced by l 2 regularization on the parameters (weights). We focus on linear neural networks as parameterizations of linear predictors. We identify the representation cost of certain sparse linear ConvNets and residual networks. In order to get a better understanding of how the architecture and parameterization affect the representation cost, we also study the reverse problem, identifying which regularizers on linear predictors (e.g., l p quasi-norms, group quasi-norms, the k-support-norm, elastic net) can be the representation cost induced by simple l 2 regularization, and designing the parameterizations that do so.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 被引用 41 次
- Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rankZihan Wang, Arthur JacotICLR 2024 · 被引用 27 次
- Feature Learning in -regularized DNNs: Attraction/Repulsion and SparsityArthur Jacot, Eugene A. Golikov, Clément Hongler, Franck GabrielNeurIPS 2022 · 被引用 22 次
- Path Regularization: A Convexity and Sparsity Inducing Regularization for Parallel ReLU NetworksTolga Ergen, Mert PilanciNeurIPS 2023 · 被引用 21 次
- Bottleneck Structure in Learned Features: Low-Dimension vs Regularity TradeoffArthur JacotNeurIPS 2023 · 被引用 20 次
它引用的顶会 Paper6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 被引用 172 次
- Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer NetworksMert Pilanci, Tolga ErgenICML 2020 · 被引用 142 次
- Revealing the Structure of Deep Neural Networks via Convex DualityTolga Ergen, Mert PilanciICML 2021 · 被引用 77 次
- Vector-output ReLU Neural Network Problems are Copositive Programs: Convex Analysis of Two Layer Networks and Polynomial-time AlgorithmsArda Sahiner, Tolga Ergen, John M. Pauly, Mert PilanciICLR 2021 · 被引用 45 次
相关 Paper
- Penalising the biases in norm regularisation enforces sparsityEtienne Boursier, Nicolas FlammarionNeurIPS 2023 · 被引用 21 次
- Topologically Densified DistributionsChristoph D. Hofer, Florian Graf, Marc Niethammer, Roland KwittICML 2020 · 被引用 15 次
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 被引用 2 次
- GULP: a prediction-based metric between representationsEnric Boix-Adserà, Hannah Lawrence, George Stepaniants, Philippe RigolletNeurIPS 2022 · 被引用 20 次
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
