On the Feature Learning in Diffusion Models
Andi Han, Wei Huang, Yuan Cao, Difan Zou
Abstract
The predominant success of diffusion models in generative modeling has spurred significant interest in understanding their theoretical foundations. In this work, we propose a feature learning framework aimed at analyzing and comparing the training dynamics of diffusion models with those of traditional classification models. Our theoretical analysis demonstrates that diffusion models, due to the denoising objective, are encouraged to learn more balanced and comprehensive representations of the data. In contrast, neural networks with a similar architecture trained for classification tend to prioritize learning specific patterns in the data, often focusing on easy-to-learn components. To support these theoretical insights, we conduct several experiments on both synthetic and real-world datasets, which empirically validate our findings and highlight the distinct feature learning dynamics in diffusion models compared to classification. PROBLEM SETTING This section introduces the problem settings for both diffusion model and classification, including the data model, neural network functions as well as training objectives and algorithm. Definition 2.1 (Data distribution). Each data sample consists of two patches, as x = [x (1)⊤ , x (2)⊤ ] ⊤ , where each patch is generated as follows: • Sample y ∈ -1, 1 uniformly with P(y = -1) = P(y = 1) = 1/2. Published as a conference paper at ICLR 2025 • Given two orthogonal signal vectors µ 1 , µ -1 , with µ 1 ⊥ µ -1 , we set x (1) = µ y , i.e., x (1) = µ 1 if y = 1 and x (1) = µ -1 if y = -1. For simplicity, we assume ∥µ 1 ∥ = ∥µ -1 ∥ = ∥µ∥. This multi-patch data model reflects the structure of image data, where each image consists of multiple patches, and only a subset of the patches are relevant to the class label, while the rest contribute as background noise. This data model has been employed in several existing studies (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f35e5a06-22b9-4d95-84a6-95dee370793fCited by top-tier papers13
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionYixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu et al.ICLR 2026 · 34 citations
- Neural Network-Based Score Estimation in Diffusion Models: Optimization and GeneralizationYinbin Han, Meisam Razaviyayn, Renyuan XuICLR 2024 · 33 citations
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingXiao Li, Zekai Zhang, Xiang Li, Siyi Chen et al.NeurIPS 2025 · 19 citations
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?Wei Huang, Andi Han, Yujin Song, Yilan Chen et al.NeurIPS 2025 · 4 citations
Builds on47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- Swing-by Dynamics in Concept Learning and Compositional GeneralizationYongyi Yang, Core Francisco Park, Ekdeep Singh Lubana, Maya Okawa et al.ICLR 2025
- Locality in Image Diffusion Models Emerges from Data StatisticsArtem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent SitzmannNeurIPS 2025 · 32 citations
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 93 citations
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 25 citations
- Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementTao Yang, Cuiling Lan, Yan Lu, Nanning ZhengNeurIPS 2024 · 41 citations
