How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
Umberto M. Tomasini, Matthieu Wyart
Abstract
Understanding what makes high-dimensional data learnable is a fundamental question in machine learning. On the one hand, it is believed that the success of deep learning lies in its ability to build a hierarchy of representations that become increasingly more abstract with depth, going from simple features like edges to more complex concepts. On the other hand, learning to be insensitive to invariances of the task, such as smooth transformations for image datasets, has been argued to be important for deep networks and it strongly correlates with their performance. In this work, we aim to explain this correlation and unify these two viewpoints. We show that by introducing sparsity to generative hierarchical models of data, the task acquires insensitivity to spatial transformations that are discrete versions of smooth transformations. In particular, we introduce the Sparse Random Hierarchy Model (SRHM), where we observe and rationalize that a hierarchical representation mirroring the hierarchical model is learnt precisely when such insensitivity is learnt, thereby explaining the strong correlation between the latter and performance. Moreover, we quantify how the sample complexity of CNNs learning the SRHM depends on both the sparsity and hierarchical structure of the task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48b34ee7-7fc0-4bb9-9c2d-74e5b58197bcCited by top-tier papers3
- Towards a theory of how the structure of language is acquired by deep neural networksFrancesco Cagnetta, Matthieu WyartNeurIPS 2024 · 33 citations
- U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical ModelsSong MeiICLR 2025
- Learning curves theory for hierarchically compositional data with power-law distributed featuresFrancesco Cagnetta, Hyunmo Kang, Matthieu WyartICML 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Practical Method for Constructing Equivariant Multilayer Perceptrons for Arbitrary Matrix GroupsMarc Finzi, Max Welling, Andrew Gordon WilsonICML 2021 · 226 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler et al.NeurIPS 2021 · 74 citations
Related papers
- Learning sparse features can lead to overfitting in neural networksLeonardo Petrini, Francesco Cagnetta, Eric Vanden-Eijnden, Matthieu WyartNeurIPS 2022 · 47 citations
- What Can Be Learnt With Wide Convolutional Neural Networks?Francesco Cagnetta, Alessandro Favero, Matthieu WyartICML 2023 · 16 citations
- Relative stability toward diffeomorphisms indicates performance in deep netsLeonardo Petrini, Alessandro Favero, Mario Geiger, Matthieu WyartNeurIPS 2021 · 16 citations
- Decoupling Semantic Similarity from Spatial Alignment for Neural NetworksTassilo Wald, Constantin Ulrich, Priyank Jaini, Gregor Köhler et al.NeurIPS 2024
- Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and CompressibilityMelih Barsbey, Lucas Prieto, Stefanos Zafeiriou, Tolga BirdalICCV 2025 · 3 citations
