Average gradient outer product as a mechanism for deep neural collapse
Daniel Beaglehole, Peter Súkeník, Marco Mondelli, Mikhail Belkin
摘要
Deep Neural Collapse (DNC) refers to the surprisingly rigid structure of the data representations in the final layers of Deep Neural Networks (DNNs). Though the phenomenon has been measured in a variety of settings, its emergence is typically explained via data-agnostic approaches, such as the unconstrained features model. In this work, we introduce a data-dependent setting where DNC forms due to feature learning through the average gradient outer product (AGOP). The AGOP is defined with respect to a learned predictor and is equal to the uncentered covariance matrix of its input-output gradients averaged over the training dataset. The Deep Recursive Feature Machine (Deep RFM) is a method that constructs a neural network by iteratively mapping the data with the AGOP and applying an untrained random feature map. We demonstrate empirically that DNC occurs in Deep RFM across standard settings as a consequence of the projection with the AGOP matrix computed at each layer. Further, we theoretically explain DNC in Deep RFM in an asymptotic setting and as a result of kernel learning. We then provide evidence that this mechanism holds for neural networks more generally. In particular, we show that the right singular vectors and values of the weights can be responsible for the majority of within-class variability collapse for DNNs trained in the feature learning regime. As observed in recent work, this singular structure is highly correlated with that of the AGOP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Linguistic Collapse: Neural Collapse in (Large) Language ModelsRobert Wu, Vardan PapyanNeurIPS 2024 · 被引用 45 次
- Compressible Dynamics in Deep Overparameterized Low-Rank Learning & AdaptationCan Yaras, Peng Wang, Laura Balzano, Qing QuICML 2024 · 被引用 29 次
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular dataDaniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan, Mikhail BelkinICLR 2026 · 被引用 18 次
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 被引用 14 次
- Neural Collapse Inspired Feature Alignment for Out-of-Distribution GeneralizationZhikang Chen, Min Zhang, Sen Cui, Haoxuan Li 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper13
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 被引用 182 次
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You 等ICML 2022 · 被引用 122 次
- Extended Unconstrained Features Model for Exploring Deep Neural CollapseTom Tirer, Joan BrunaICML 2022 · 被引用 118 次
相关 Paper
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 被引用 12 次
- The Prevalence of Neural Collapse in Neural Multivariate RegressionGeorge Andriopoulos, Zixuan Dong, Li Guo, Zifan Zhao 等NeurIPS 2024 · 被引用 24 次
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
- Deep Neural Collapse Is Provably Optimal for the Deep Unconstrained Features ModelPeter Súkeník, Marco Mondelli, Christoph H. LampertNeurIPS 2023 · 被引用 51 次
- The Persistence of Neural Collapse Despite Low-Rank BiasConnall Garrod, Jonathan P. KeatingNeurIPS 2025 · 被引用 2 次
