Probing the Latent Hierarchical Structure of Data via Diffusion Models
Antonio Sclocchi, Alessandro Favero, Noam Itzhak Levi, Matthieu Wyart
摘要
High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce. Likewise, accessing the latent variables underlying such a data structure remains a challenge. In this work, we show that forward-backward experiments in diffusion-based models, where data is noised and then denoised to generate new samples, are a promising tool to probe the latent structure of data. We predict in simple hierarchical models that, in this process, changes in data occur by correlated chunks, with a length scale that diverges at a noise level where a phase transition is known to take place. Remarkably, we confirm this prediction in both text and image datasets using state-of-the-art diffusion models. Our results show how latent variable changes manifest in the data and establish how to measure these effects in real data using diffusion models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A solvable model of learning generative diffusion: theory and insightsHugo Cui, Cengiz Pehlevan, Yue M. LuNeurIPS 2025 · 被引用 11 次
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 被引用 10 次
- The Computational Advantage of Depth in Learning High-Dimensional Hierarchical TargetsYatin Dandi, Luca Pesce, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 被引用 3 次
- Biased Generalization in Diffusion ModelsLuca Saglietti, Luca Biggio, Jerome Garnier-Brun, Davide Beltrame 等ICML 2026 · 被引用 2 次
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 被引用 1 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- How Compositional Generalization and Creativity Improve as Diffusion Models are TrainedAlessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard 等ICML 2025
- Interpretable Diffusion via Information DecompositionXianghao Kong, Ollie Liu, Han Li, Dani Yogatama 等ICLR 2024 · 被引用 37 次
- The Entropic Signature of Class Speciation in Diffusion ModelsFlorian Handke, Dejan Stancevic, Felix Koulischer, Thomas Demeester 等ICML 2026 · 被引用 5 次
- How Diffusion Models Learn to Factorize and ComposeQiyao Liang, Ziming Liu, Mitchell Ostrow, Ila FieteNeurIPS 2024 · 被引用 17 次
- Truncated Diffusion Probabilistic Models and Diffusion-based Adversarial Auto-EncodersHuangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan ZhouICLR 2023 · 被引用 19 次
