Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement
Tao Yang, Cuiling Lan, Yan Lu, Nanning Zheng
摘要
Disentangled representation learning strives to extract the intrinsic factors within observed data. Factorizing these representations in an unsupervised manner is notably challenging and usually requires tailored loss functions or specific structural designs. In this paper, we introduce a new perspective and framework, demonstrating that diffusion models with cross-attention can serve as a powerful inductive bias to facilitate the learning of disentangled representations. We propose to encode an image to a set of concept tokens and treat them as the condition of the latent diffusion for image reconstruction, where cross-attention over the concept tokens is used to bridge the interaction between the encoder and diffusion. Without any additional regularization, this framework achieves superior disentanglement performance on the benchmark datasets, surpassing all previous methods with intricate designs. We have conducted comprehensive ablation studies and visualization analysis, shedding light on the functioning of this model. This is the first work to reveal the potent disentanglement capability of diffusion models with cross-attention, requiring no complex designs. We anticipate that our findings will inspire more investigation on exploring diffusion for disentangled representation learning towards more sophisticated data analysis and understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang 等CVPR 2026 · 被引用 46 次
- Causality-Guided Prompt Learning for Vision-Language Models via Visual GranulationMengyu Gao, Qiulei DongICCV 2025 · 被引用 2 次
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual GenerationLei Tong, Zhihua Liu, Chaochao Lu, Dino Oglic 等ICML 2026 · 被引用 2 次
- DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across ModalitiesHedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn 等ICLR 2026 · 被引用 2 次
- Understanding Implosion in Text-to-Image Generative ModelsWenxin Ding, Cathy Yuanchen Li, Shawn Shan, Ben Y. Zhao 等CCS 2024 · 被引用 2 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 被引用 1,049 次
相关 Paper
- Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation LearningAncong Wu, Wei-Shi ZhengAAAI 2024 · 被引用 10 次
- DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic ModelsTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2023 · 被引用 41 次
- Visual Concepts TokenizationTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2022 · 被引用 19 次
- Exploring Diffusion Time-steps for Unsupervised Representation LearningZhongqi Yue, Jiankun Wang, Qianru Sun, Lei Ji 等ICLR 2024 · 被引用 33 次
- InfoDiffusion: Representation Learning Using Information Maximizing Diffusion ModelsYingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan 等ICML 2023 · 被引用 64 次
