Hierarchical nucleation in deep neural networks
Diego Doimo, Aldo Glielmo, Alessio Ansuini, Alessandro Laio
摘要
Deep convolutional networks (DCNs) learn meaningful representations where data that share the same abstract characteristics are positioned closer and closer. Understanding these representations and how they are generated is of unquestioned practical and theoretical interest. In this work we study the evolution of the probability density of the ImageNet dataset across the hidden layers in some state-of-the-art DCNs. We find that the initial layers generate a unimodal probability density getting rid of any structure irrelevant for classification. In subsequent layers density peaks arise in a hierarchical fashion that mirrors the semantic hierarchy of the concepts. Density peaks corresponding to single categories appear only close to the output and via a very sharp transition which resembles the nucleation process of a heterogeneous liquid. This process leaves a footprint in the probability density of the output layer where the topography of the peaks allows reconstructing the semantic relationships of the categories.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- The geometry of hidden representations of large transformer modelsLucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio 等NeurIPS 2023 · 被引用 148 次
- Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian ManifoldCan Yaras, Peng Wang, Zhihui Zhu, Laura Balzano 等NeurIPS 2022 · 被引用 60 次
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 被引用 58 次
- The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language ModelsDiego Doimo, Alessandro Serra, Alessio Ansuini, Alberto CazzanigaNeurIPS 2024 · 被引用 21 次
- Beyond Scalars: Concept-Based Alignment Analysis in Vision TransformersJohanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek 等NeurIPS 2025 · 被引用 11 次
相关 Paper
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 被引用 78 次
- Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant FrequenciesHannah Pinson, Joeri Lenaerts, Vincent GinisICML 2023 · 被引用 8 次
- What do CNNs Learn in the First Layer and Why? A Linear Systems PerspectiveRhea Chowers, Yair WeissICML 2023 · 被引用 5 次
- Finding Representative Interpretations on Convolutional Neural NetworksPeter Cho-Ho Lam, Lingyang Chu, Maxim Torgonskiy, Jian Pei 等ICCV 2021 · 被引用 7 次
- Prune and distill: similar reformatting of image information along rat visual cortex and deep neural networksPaolo Muratore, Sina Tafazoli, Eugenio Piasini, Alessandro Laio 等NeurIPS 2022 · 被引用 11 次
