Probing the Latent Hierarchical Structure of Data via Diffusion Models
Antonio Sclocchi, Alessandro Favero, Noam Itzhak Levi, Matthieu Wyart
Abstract
High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce. Likewise, accessing the latent variables underlying such a data structure remains a challenge. In this work, we show that forward-backward experiments in diffusion-based models, where data is noised and then denoised to generate new samples, are a promising tool to probe the latent structure of data. We predict in simple hierarchical models that, in this process, changes in data occur by correlated chunks, with a length scale that diverges at a noise level where a phase transition is known to take place. Remarkably, we confirm this prediction in both text and image datasets using state-of-the-art diffusion models. Our results show how latent variable changes manifest in the data and establish how to measure these effects in real data using diffusion models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- A solvable model of learning generative diffusion: theory and insightsHugo Cui, Cengiz Pehlevan, Yue M. LuNeurIPS 2025 · 11 citations
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 10 citations
- The Computational Advantage of Depth in Learning High-Dimensional Hierarchical TargetsYatin Dandi, Luca Pesce, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 3 citations
- Biased Generalization in Diffusion ModelsLuca Saglietti, Luca Biggio, Jerome Garnier-Brun, Davide Beltrame et al.ICML 2026 · 2 citations
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- How Compositional Generalization and Creativity Improve as Diffusion Models are TrainedAlessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard et al.ICML 2025
- Interpretable Diffusion via Information DecompositionXianghao Kong, Ollie Liu, Han Li, Dani Yogatama et al.ICLR 2024 · 37 citations
- The Entropic Signature of Class Speciation in Diffusion ModelsFlorian Handke, Dejan Stancevic, Felix Koulischer, Thomas Demeester et al.ICML 2026 · 5 citations
- How Diffusion Models Learn to Factorize and ComposeQiyao Liang, Ziming Liu, Mitchell Ostrow, Ila FieteNeurIPS 2024 · 17 citations
- Truncated Diffusion Probabilistic Models and Diffusion-based Adversarial Auto-EncodersHuangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan ZhouICLR 2023 · 19 citations
