Hierarchical nucleation in deep neural networks
Diego Doimo, Aldo Glielmo, Alessio Ansuini, Alessandro Laio
Abstract
Deep convolutional networks (DCNs) learn meaningful representations where data that share the same abstract characteristics are positioned closer and closer. Understanding these representations and how they are generated is of unquestioned practical and theoretical interest. In this work we study the evolution of the probability density of the ImageNet dataset across the hidden layers in some state-of-the-art DCNs. We find that the initial layers generate a unimodal probability density getting rid of any structure irrelevant for classification. In subsequent layers density peaks arise in a hierarchical fashion that mirrors the semantic hierarchy of the concepts. Density peaks corresponding to single categories appear only close to the output and via a very sharp transition which resembles the nucleation process of a heterogeneous liquid. This process leaves a footprint in the probability density of the output layer where the topography of the peaks allows reconstructing the semantic relationships of the categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09d5a504-ed5d-4dee-a5cd-190e69d113adCited by top-tier papers15
- The geometry of hidden representations of large transformer modelsLucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio et al.NeurIPS 2023 · 148 citations
- Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian ManifoldCan Yaras, Peng Wang, Zhihui Zhu, Laura Balzano et al.NeurIPS 2022 · 60 citations
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 58 citations
- The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language ModelsDiego Doimo, Alessandro Serra, Alessio Ansuini, Alberto CazzanigaNeurIPS 2024 · 21 citations
- Beyond Scalars: Concept-Based Alignment Analysis in Vision TransformersJohanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek et al.NeurIPS 2025 · 11 citations
Related papers
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 78 citations
- Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant FrequenciesHannah Pinson, Joeri Lenaerts, Vincent GinisICML 2023 · 8 citations
- What do CNNs Learn in the First Layer and Why? A Linear Systems PerspectiveRhea Chowers, Yair WeissICML 2023 · 5 citations
- Finding Representative Interpretations on Convolutional Neural NetworksPeter Cho-Ho Lam, Lingyang Chu, Maxim Torgonskiy, Jian Pei et al.ICCV 2021 · 7 citations
- Prune and distill: similar reformatting of image information along rat visual cortex and deep neural networksPaolo Muratore, Sina Tafazoli, Eugenio Piasini, Alessandro Laio et al.NeurIPS 2022 · 11 citations
