Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
X. Y. Han, Vardan Papyan, David L. Donoho
摘要
The recently discovered Neural Collapse (NC) phenomenon occurs pervasively in today's deep net training paradigm of driving cross-entropy (CE) loss towards zero. During NC, last-layer features collapse to their class-means, both classifiers and class-means collapse to the same Simplex Equiangular Tight Frame, and classifier behavior collapses to the nearest-class-mean decision rule. Recent works demonstrated that deep nets trained with mean squared error (MSE) loss perform comparably to those trained with CE. As a preliminary, we empirically establish that NC emerges in such MSE-trained deep nets as well through experiments on three canonical networks and five benchmark datasets. We provide, in a Google Colab notebook, PyTorch code for reproducing MSE-NC and CE-NC: at https://colab.research.google.com/github/neuralcollapse/neuralcollapse/blob/main/neuralcollapse.ipynb. The analytically-tractable MSE loss offers more mathematical opportunities than the hard-to-analyze CE loss, inspiring us to leverage MSE loss towards the theoretical investigation of NC. We develop three main contributions: (I) We show a new decomposition of the MSE loss into (A) terms directly interpretable through the lens of NC and which assume the last-layer classifier is exactly the least-squares classifier; and (B) a term capturing the deviation from this least-squares classifier. (II) We exhibit experiments on canonical datasets and networks demonstrating that term-(B) is negligible during training. This motivates us to introduce a new theoretical construct: the central path, where the linear classifier stays MSE-optimal for feature activations throughout the dynamics. (III) By studying renormalized gradient flow along the central path, we derive exact dynamics that predict NC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper92
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 被引用 152 次
- Inducing Neural Collapse in Imbalanced Learning: Do We Really Need a Learnable Classifier at the End of Deep Neural Network?Yibo Yang, Shixiang Chen, Xiangtai Li, Liang Xie 等NeurIPS 2022 · 被引用 144 次
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You 等ICML 2022 · 被引用 122 次
- Extended Unconstrained Features Model for Exploring Deep Neural CollapseTom Tirer, Joan BrunaICML 2022 · 被引用 118 次
- On the Role of Neural Collapse in Transfer LearningTomer Galanti, András György, Marcus HutterICLR 2022 · 被引用 114 次
它引用的顶会 Paper4
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 等NeurIPS 2021 · 被引用 303 次
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- Revealing the Structure of Deep Neural Networks via Convex DualityTolga Ergen, Mert PilanciICML 2021 · 被引用 77 次
相关 Paper
- Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced DataHien Dang, Tho Tran Huu, Stanley J. Osher, Hung Tran-The 等ICML 2023 · 被引用 44 次
- Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features ModelHien Dang, Tho Tran Huu, Tan Minh Nguyen, Nhat HoICML 2024 · 被引用 19 次
- Neural (Tangent Kernel) CollapseMariia Seleznova, Dana Weitzner, Raja Giryes, Gitta Kutyniok 等NeurIPS 2023 · 被引用 23 次
- Generalizing and Decoupling Neural Collapse via Hyperspherical Uniformity GapWeiyang Liu, Longhui Yu, Adrian Weller, Bernhard SchölkopfICLR 2023 · 被引用 3 次
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu 等NeurIPS 2022 · 被引用 93 次
