Feature Learning in -regularized DNNs: Attraction/Repulsion and Sparsity
Arthur Jacot, Eugene A. Golikov, Clément Hongler, Franck Gabriel
Abstract
We study the loss surface of DNNs with regularization. We show that the loss in terms of the parameters can be reformulated into a loss in terms of the layerwise activations of the training set. This reformulation reveals the dynamics behind feature learning: each hidden representations are optimal w.r.t. to an attraction/repulsion problem and interpolate between the input and output representations, keeping as little information from the input as necessary to construct the activation of the next layer. For positively homogeneous non-linearities, the loss can be further reformulated in terms of the covariances of the hidden representations, which takes the form of a partially convex optimization over a convex cone. This second reformulation allows us to prove a sparsity result for homogeneous DNNs: any local minimum of the -regularized loss can be achieved with at most neurons in each hidden layer (where is the size of the training set). We show that this bound is tight by giving an example of a local minimum that requires hidden neurons. But we also observe numerically that in more traditional settings much less than neurons are required to reach the minima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb5c38b2-d72d-4b1e-8971-0bc36f453037Cited by top-tier papers12
- On the Stepwise Nature of Self-Supervised LearningJames B. Simon, Maksis Knutins, Liu Ziyin, Daniel Geisz et al.ICML 2023 · 45 citations
- Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rankZihan Wang, Arthur JacotICLR 2024 · 27 citations
- An exactly solvable model for emergence and scaling laws in the multitask sparse parity problemYoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard et al.NeurIPS 2024 · 20 citations
- Bottleneck Structure in Learned Features: Low-Dimension vs Regularity TradeoffArthur JacotNeurIPS 2023 · 20 citations
- Mixed Dynamics In Linear Networks: Unifying the Lazy and Active RegimesZhenfeng Tu, Santiago Aranguri, Arthur JacotNeurIPS 2024 · 18 citations
Builds on5
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Representation Costs of Linear Neural Networks: Analysis and DesignZhen Dai, Mina Karzand, Nathan SrebroNeurIPS 2021 · 34 citations
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter et al.ICLR 2021 · 22 citations
Related papers
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 2 citations
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 33 citations
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 1 citation
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
- The Persistence of Neural Collapse Despite Low-Rank BiasConnall Garrod, Jonathan P. KeatingNeurIPS 2025 · 2 citations
