Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
Arthur Jacot, Peter Súkeník, Zihan Wang, Marco Mondelli
Abstract
Deep neural networks (DNNs) at convergence consistently represent the training data in the last layer via a geometric structure referred to as neural collapse. This empirical evidence has spurred a line of theoretical research aimed at proving the emergence of neural collapse, mostly focusing on the unconstrained features model. Here, the features of the penultimate layer are free variables, which makes the model data-agnostic and puts into question its ability to capture DNN training. Our work addresses the issue, moving away from unconstrained features and studying DNNs that end with at least two linear layers. We first prove generic guarantees on neural collapse that assume (i) low training error and balancedness of linear layers (for within-class variability collapse), and (ii) bounded conditioning of the features before the linear part (for orthogonality of class-means, and their alignment with weight matrices). The balancedness refers to the fact that W ⊤ ℓ+1 W ℓ+1 ≈ W ℓ W ⊤ ℓ for any pair of consecutive weight matrices of the linear part, and the bounded conditioning requires a well-behaved ratio between largest and smallest non-zero singular values of the features. We then show that such assumptions hold for gradient descent training with weight decay: (i) for networks with a wide first layer, we prove low training error and balancedness, and (ii) for solutions that are either nearly optimal or stable under large learning rates, we additionally prove the bounded conditioning. Taken together, our results are the first to show neural collapse in the end-to-end training of DNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a018b07-3a23-4381-8a49-f29cff9fc593Cited by top-tier papers11
- Neural Collapse is Globally Optimal in Deep Regularized ResNets and TransformersPeter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2025 · 12 citations
- Explaining Grokking and Information Bottleneck through Neural Collapse EmergenceKeitaro Sakamoto, Issei SatoICLR 2026 · 5 citations
- Weight Decay Improves Language Model PlasticityTessa Han, Sebastian Bordt, Hanlin Zhang, Sham KakadeICML 2026 · 3 citations
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataHancheng Min, Zhihui Zhu, René VidalNeurIPS 2025 · 3 citations
- Heads collapse, features stay: Why Replay needs big buffersGiulia Lanzillotta, Damiano Meier, Thomas HofmannICLR 2026 · 3 citations
Builds on29
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 182 citations
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You et al.ICML 2022 · 122 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
- Extended Unconstrained Features Model for Exploring Deep Neural CollapseTom Tirer, Joan BrunaICML 2022 · 118 citations
Related papers
- Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field RegimeDiyuan Wu, Marco MondelliICML 2025
- Deep Neural Collapse Is Provably Optimal for the Deep Unconstrained Features ModelPeter Súkeník, Marco Mondelli, Christoph H. LampertNeurIPS 2023 · 51 citations
- Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced DataHien Dang, Tho Tran Huu, Stanley J. Osher, Hung Tran-The et al.ICML 2023 · 44 citations
- The Implicit Bias of Depth: From Neural Collapse to Softmax CodesConnall Garrod, Jonathan Keating, Christos ThrampoulidisICML 2026
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 14 citations
