Representation Costs of Linear Neural Networks: Analysis and Design
Zhen Dai, Mina Karzand, Nathan Srebro
Abstract
For different parameterizations (mappings from parameters to predictors), we study the regularization cost in predictor space induced by l 2 regularization on the parameters (weights). We focus on linear neural networks as parameterizations of linear predictors. We identify the representation cost of certain sparse linear ConvNets and residual networks. In order to get a better understanding of how the architecture and parameterization affect the representation cost, we also study the reverse problem, identifying which regularizers on linear predictors (e.g., l p quasi-norms, group quasi-norms, the k-support-norm, elastic net) can be the representation cost induced by simple l 2 regularization, and designing the parameterizations that do so.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3401a40-d390-469f-a7c3-4b6f148b87d6Cited by top-tier papers20
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
- Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rankZihan Wang, Arthur JacotICLR 2024 · 27 citations
- Feature Learning in -regularized DNNs: Attraction/Repulsion and SparsityArthur Jacot, Eugene A. Golikov, Clément Hongler, Franck GabrielNeurIPS 2022 · 22 citations
- Path Regularization: A Convexity and Sparsity Inducing Regularization for Parallel ReLU NetworksTolga Ergen, Mert PilanciNeurIPS 2023 · 21 citations
- Bottleneck Structure in Learned Features: Low-Dimension vs Regularity TradeoffArthur JacotNeurIPS 2023 · 20 citations
Builds on6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer NetworksMert Pilanci, Tolga ErgenICML 2020 · 142 citations
- Revealing the Structure of Deep Neural Networks via Convex DualityTolga Ergen, Mert PilanciICML 2021 · 77 citations
- Vector-output ReLU Neural Network Problems are Copositive Programs: Convex Analysis of Two Layer Networks and Polynomial-time AlgorithmsArda Sahiner, Tolga Ergen, John M. Pauly, Mert PilanciICLR 2021 · 45 citations
Related papers
- Penalising the biases in norm regularisation enforces sparsityEtienne Boursier, Nicolas FlammarionNeurIPS 2023 · 21 citations
- Topologically Densified DistributionsChristoph D. Hofer, Florian Graf, Marc Niethammer, Roland KwittICML 2020 · 15 citations
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 2 citations
- GULP: a prediction-based metric between representationsEnric Boix-Adserà, Hannah Lawrence, George Stepaniants, Philippe RigolletNeurIPS 2022 · 20 citations
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
