The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
Michael Munn, Benoit Dherin, Javier Gonzalvo
Abstract
Many of the recent remarkable advances in computer vision and language models can be attributed to the success of transfer learning via the pre-training of large foundation models. However, a theoretical framework which explains this empirical success is incomplete and remains an active area of research. Flatness of the loss surface and neural collapse have recently emerged as useful pre-training metrics which shed light on the implicit biases underlying pre-training. In this paper, we explore the geometric complexity of a model's learned representations as a fundamental mechanism that relates these two concepts. We show through experiments and theory that mechanisms which affect the geometric complexity of the pre-trained network also influence the neural collapse. Furthermore, we show how this effect of the geometric complexity generalizes to the neural collapse of new classes as well, thus encouraging better performance on downstream tasks, particularly in the few-shot setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08e67425-2e9d-4b3b-b8dd-dc9f1a0d88ebCited by top-tier papers5
- Neural Collapse in Multi-Task LearningYoujun Wang, Boqi Li, Xin Zou, Weiwei LiuICLR 2026 · 16 citations
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek et al.NeurIPS 2025 · 9 citations
- A Bayesian Model Selection Criterion for Selecting Pretraining CheckpointsMichael Munn, Susan WeiICML 2025
- Your Dissimilarities Define You: Complementary Learning Exploiting Class DiversitiesDimitrios Katsikas, Nikolaos Passalis, Anastasios TefasCVPR 2026
- On the Local Complexity of Linear Regions in Deep ReLU NetworksNiket Patel, Guido MontúfarICML 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
Related papers
- On the Role of Neural Collapse in Transfer LearningTomer Galanti, András György, Marcus HutterICLR 2022 · 114 citations
- How Far Pre-trained Models Are from Neural Collapse on the Target Dataset Informs their TransferabilityZijian Wang, Yadan Luo, Liang Zheng, Zi Huang et al.ICCV 2023 · 33 citations
- Quantifying the Variability Collapse of Neural NetworksJing Xu, Haoxiong LiuICML 2023 · 10 citations
- Explaining Grokking and Information Bottleneck through Neural Collapse EmergenceKeitaro Sakamoto, Issei SatoICLR 2026 · 5 citations
- Frustratingly Easy Transferability EstimationLong-Kai Huang, Junzhou Huang, Yu Rong, Qiang Yang et al.ICML 2022 · 71 citations
