The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
Michael Munn, Benoit Dherin, Javier Gonzalvo
摘要
Many of the recent remarkable advances in computer vision and language models can be attributed to the success of transfer learning via the pre-training of large foundation models. However, a theoretical framework which explains this empirical success is incomplete and remains an active area of research. Flatness of the loss surface and neural collapse have recently emerged as useful pre-training metrics which shed light on the implicit biases underlying pre-training. In this paper, we explore the geometric complexity of a model's learned representations as a fundamental mechanism that relates these two concepts. We show through experiments and theory that mechanisms which affect the geometric complexity of the pre-trained network also influence the neural collapse. Furthermore, we show how this effect of the geometric complexity generalizes to the neural collapse of new classes as well, thus encouraging better performance on downstream tasks, particularly in the few-shot setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Neural Collapse in Multi-Task LearningYoujun Wang, Boqi Li, Xin Zou, Weiwei LiuICLR 2026 · 被引用 16 次
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek 等NeurIPS 2025 · 被引用 9 次
- A Bayesian Model Selection Criterion for Selecting Pretraining CheckpointsMichael Munn, Susan WeiICML 2025
- Your Dissimilarities Define You: Complementary Learning Exploiting Class DiversitiesDimitrios Katsikas, Nikolaos Passalis, Anastasios TefasCVPR 2026
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
相关 Paper
- On the Role of Neural Collapse in Transfer LearningTomer Galanti, András György, Marcus HutterICLR 2022 · 被引用 114 次
- How Far Pre-trained Models Are from Neural Collapse on the Target Dataset Informs their TransferabilityZijian Wang, Yadan Luo, Liang Zheng, Zi Huang 等ICCV 2023 · 被引用 33 次
- Quantifying the Variability Collapse of Neural NetworksJing Xu, Haoxiong LiuICML 2023 · 被引用 10 次
- Explaining Grokking and Information Bottleneck through Neural Collapse EmergenceKeitaro Sakamoto, Issei SatoICLR 2026 · 被引用 5 次
- Frustratingly Easy Transferability EstimationLong-Kai Huang, Junzhou Huang, Yu Rong, Qiang Yang 等ICML 2022 · 被引用 71 次
