On the Feature Learning in Diffusion Models
Andi Han, Wei Huang, Yuan Cao, Difan Zou
摘要
The predominant success of diffusion models in generative modeling has spurred significant interest in understanding their theoretical foundations. In this work, we propose a feature learning framework aimed at analyzing and comparing the training dynamics of diffusion models with those of traditional classification models. Our theoretical analysis demonstrates that diffusion models, due to the denoising objective, are encouraged to learn more balanced and comprehensive representations of the data. In contrast, neural networks with a similar architecture trained for classification tend to prioritize learning specific patterns in the data, often focusing on easy-to-learn components. To support these theoretical insights, we conduct several experiments on both synthetic and real-world datasets, which empirically validate our findings and highlight the distinct feature learning dynamics in diffusion models compared to classification. PROBLEM SETTING This section introduces the problem settings for both diffusion model and classification, including the data model, neural network functions as well as training objectives and algorithm. Definition 2.1 (Data distribution). Each data sample consists of two patches, as x = [x (1)⊤ , x (2)⊤ ] ⊤ , where each patch is generated as follows: • Sample y ∈ -1, 1 uniformly with P(y = -1) = P(y = 1) = 1/2. Published as a conference paper at ICLR 2025 • Given two orthogonal signal vectors µ 1 , µ -1 , with µ 1 ⊥ µ -1 , we set x (1) = µ y , i.e., x (1) = µ 1 if y = 1 and x (1) = µ -1 if y = -1. For simplicity, we assume ∥µ 1 ∥ = ∥µ -1 ∥ = ∥µ∥. This multi-patch data model reflects the structure of image data, where each image consists of multiple patches, and only a subset of the patches are relevant to the class label, while the rest contribute as background noise. This data model has been employed in several existing studies (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionYixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu 等ICLR 2026 · 被引用 34 次
- Neural Network-Based Score Estimation in Diffusion Models: Optimization and GeneralizationYinbin Han, Meisam Razaviyayn, Renyuan XuICLR 2024 · 被引用 33 次
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingXiao Li, Zekai Zhang, Xiang Li, Siyi Chen 等NeurIPS 2025 · 被引用 19 次
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi 等ICLR 2026 · 被引用 14 次
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?Wei Huang, Andi Han, Yujin Song, Yilan Chen 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
相关 Paper
- Swing-by Dynamics in Concept Learning and Compositional GeneralizationYongyi Yang, Core Francisco Park, Ekdeep Singh Lubana, Maya Okawa 等ICLR 2025
- Locality in Image Diffusion Models Emerges from Data StatisticsArtem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent SitzmannNeurIPS 2025 · 被引用 32 次
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 被引用 93 次
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 被引用 25 次
- Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementTao Yang, Cuiling Lan, Yan Lu, Nanning ZhengNeurIPS 2024 · 被引用 41 次
