How Diffusion Models Learn to Factorize and Compose
Qiyao Liang, Ziming Liu, Mitchell Ostrow, Ila Fiete
Abstract
Diffusion models are capable of generating photo-realistic images that combine elements which likely do not appear together in the training set, demonstrating the ability to compositionally generalize. Nonetheless, the precise mechanism of compositionality and how it is acquired through training remains elusive. Inspired by cognitive neuroscientific approaches, we consider a highly reduced setting to examine whether and when diffusion models learn semantically meaningful and factorized representations of composable features. We performed extensive controlled experiments on conditional Denoising Diffusion Probabilistic Models (DDPMs) trained to generate various forms of 2D Gaussian bump images. We found that the models learn factorized but not fully continuous manifold representations for encoding continuous features of variation underlying the data. With such representations, models demonstrate superior feature compositionality but limited ability to interpolate over unseen values of a given feature. Our experimental results further demonstrate that diffusion models can attain compositionality with few compositional examples, suggesting a more efficient way to train DDPMs. Finally, we connect manifold formation in diffusion models to percolation theory in physics, offering insight into the sudden onset of factorized representation learning. Our thorough toy experiments thus contribute a deeper understanding of how diffusion models capture compositional structure in data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36af31a8-67de-403e-86fe-3e940e008c0cCited by top-tier papers5
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 11 citations
- Compositional Generalization via Forced Rendering of Disentangled LatentsQiyao Liang, Daoyuan Qian, Liu Ziyin, Ila R. FieteICML 2025
- CoInD: Enabling Logical Compositions in Diffusion ModelsSachit Gaudi, Gautam Sreekumar, Vishnu BoddetiICLR 2025
- Concept Lancet: Image Editing with Compositional Representation TransplantJinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Hancheng Min et al.CVPR 2025
- Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM GuidanceDongmin Park, Sebin Kim, Taehong Moon, Minkyu Kim et al.ICLR 2025
Builds on8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 113 citations
- The role of Disentanglement in GeneralisationMilton Llera Montero, Casimir J. H. Ludwig, Rui Ponte Costa, Gaurav Malhotra et al.ICLR 2021 · 97 citations
- Unsupervised Model Selection for Variational Disentangled Representation LearningSunny Duan, Loic Matthey, Andre Saraiva, Nick Watters et al.ICLR 2020 · 87 citations
- Compositional Generalization from First PrinciplesThaddäus Wiedemer, Prasanna Mayilvahanan, Matthias Bethge, Wieland BrendelNeurIPS 2023 · 78 citations
Related papers
- How Compositional Generalization and Creativity Improve as Diffusion Models are TrainedAlessandro Favero, Antonio Sclocchi, Francesco Cagnetta, Pascal Frossard et al.ICML 2025
- Going beyond Compositions, DDPMs Can Produce Zero-Shot InterpolationsJustin Deschenaux, Igor Krawczuk, Grigorios Chrysos, Volkan CevherICML 2024 · 6 citations
- DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic ModelsTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2023 · 41 citations
- Composer: Creative and Controllable Image Synthesis with Composable ConditionsLianghua Huang, Di Chen, Yu Liu, Yujun Shen et al.ICML 2023 · 371 citations
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 276 citations
