Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation Learning
Ancong Wu, Wei-Shi Zheng
Abstract
Unsupervised disentangled representation learning aims to recover semantically meaningful factors from real-world data without supervision, which is significant for model generalization and interpretability. Current methods mainly rely on assumptions of independence or informativeness of factors, regardless of interpretability. Intuitively, visually interpretable concepts better align with human-defined factors. However, exploiting visual interpretability as inductive bias is still under-explored. Inspired by the observation that most explanatory image factors can be represented by ``content + mask'', we propose a content-mask factorization network (CMFNet) to decompose an image into different groups of content codes and masks, which are further combined as content masks to represent different visual concepts. To ensure informativeness of the representations, the CMFNet is jointly learned with a generator conditioned on the content masks for reconstructing the input image. The conditional generator employs a diffusion model to leverage its robust distribution modeling capability. Our model is called the Factorized Diffusion Autoencoder (FDAE). To enhance disentanglement of visual concepts, we propose a content decorrelation loss and a mask entropy loss to decorrelate content masks in latent space and spatial space, respectively. Experiments on Shapes3d, MPI3D and Cars3d show that our method achieves advanced performance and can generate visually interpretable concept-specific masks. Source code and supplementary materials are available at https://github.com/wuancong/FDAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 933043c5-f6c7-4e06-9e9d-b8fa3cfbde17Cited by top-tier papers4
- Unsupervised Region-Based Image Editing of Denoising Diffusion ModelsZixiang Li, Yue Song, Renshuai Tao, Xiaohong Jia et al.AAAI 2025 · 1 citation
- ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation LearningGe Gao, Di Xiong, Zeke Xie, Jian Yang et al.ICML 2026
- CLOAK: Contrastive Guidance for Latent Diffusion-Based Data ObfuscationXin Yang, Omid ArdakanianUbiComp 2026
- Diffusion Bridge AutoEncoders for Unsupervised Representation LearningYeongmin Kim, Kwanghyeon Lee, Minsang Park, Byeonghu Na et al.ICLR 2025
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
Related papers
- Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementTao Yang, Cuiling Lan, Yan Lu, Nanning ZhengNeurIPS 2024 · 41 citations
- DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic ModelsTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2023 · 41 citations
- The Hidden Language of Diffusion ModelsHila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin et al.ICLR 2024 · 38 citations
- InfoDiffusion: Representation Learning Using Information Maximizing Diffusion ModelsYingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan et al.ICML 2023 · 64 citations
- Where and What? Examining Interpretable Disentangled RepresentationsXinqi Zhu, Chang Xu, Dacheng TaoCVPR 2021
