Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin, Wenjun Zeng
Abstract
Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality, these factors may exhibit correlations, which off-the-shelf solutions have yet to properly address. To tackle this challenge, we introduce a bidirectional weighted graph-based framework, to learn factorized attributes and their interrelations within complex data. Specifically, we propose a -VAE based module to extract factors as the initial nodes of the graph, and leverage the multimodal large language model (MLLM) to discover and rank latent correlations, thereby updating the weighted edges. By integrating these complementary modules, our model successfully achieves fine-grained, practical and unsupervised disentanglement. Experiments demonstrate our method's superior performance in disentanglement and reconstruction. Furthermore, the model inherits enhanced interpretability and generalizability from MLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d4581a0-6701-45d6-a089-4d51573f8979Cited by top-tier papers3
- Coordinated Disentanglement with Iterative Mode Discovery Under Hidden CorrelationsRong Hu, Ling ChenICML 2026
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningQi Wang, Zhipeng Zhang, Baao Xie, Xin Jin et al.ICCV 2025
- Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning FrameworkJinrong Zhang, Zhaoyang Xu, Xusheng He, Xinrui Li et al.CVPR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 1,001 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
Related papers
- Revealing Multimodal Causality with Large Language ModelsJin Li, Shoujin Wang, Qi Zhang, Feng Liu et al.NeurIPS 2025 · 5 citations
- Interpretable Deep Graph Generation with Node-edge Co-disentanglementXiaojie Guo, Liang Zhao, Zhao Qin, Lingfei Wu et al.KDD 2020 · 28 citations
- Attribute-driven Disentangled Representation Learning for Multimodal RecommendationZhenyang Li, Fan Liu, Yinwei Wei, Zhiyong Cheng et al.ACM MM 2024 · 17 citations
- CausalVAE: Disentangled Representation Learning via Neural Structural Causal ModelsMengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen et al.CVPR 2021
- Towards Building A Group-based Unsupervised Representation Disentanglement FrameworkTao Yang, Xuanchi Ren, Yuwang Wang, Wenjun Zeng et al.ICLR 2022 · 36 citations
