Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin, Wenjun Zeng
摘要
Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality, these factors may exhibit correlations, which off-the-shelf solutions have yet to properly address. To tackle this challenge, we introduce a bidirectional weighted graph-based framework, to learn factorized attributes and their interrelations within complex data. Specifically, we propose a -VAE based module to extract factors as the initial nodes of the graph, and leverage the multimodal large language model (MLLM) to discover and rank latent correlations, thereby updating the weighted edges. By integrating these complementary modules, our model successfully achieves fine-grained, practical and unsupervised disentanglement. Experiments demonstrate our method's superior performance in disentanglement and reconstruction. Furthermore, the model inherits enhanced interpretability and generalizability from MLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Coordinated Disentanglement with Iterative Mode Discovery Under Hidden CorrelationsRong Hu, Ling ChenICML 2026
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningQi Wang, Zhipeng Zhang, Baao Xie, Xin Jin 等ICCV 2025
- Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning FrameworkJinrong Zhang, Zhaoyang Xu, Xusheng He, Xinrui Li 等CVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 被引用 1,001 次
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji 等ICML 2024 · 被引用 786 次
相关 Paper
- Revealing Multimodal Causality with Large Language ModelsJin Li, Shoujin Wang, Qi Zhang, Feng Liu 等NeurIPS 2025 · 被引用 5 次
- Interpretable Deep Graph Generation with Node-edge Co-disentanglementXiaojie Guo, Liang Zhao, Zhao Qin, Lingfei Wu 等KDD 2020 · 被引用 28 次
- Attribute-driven Disentangled Representation Learning for Multimodal RecommendationZhenyang Li, Fan Liu, Yinwei Wei, Zhiyong Cheng 等ACM MM 2024 · 被引用 17 次
- CausalVAE: Disentangled Representation Learning via Neural Structural Causal ModelsMengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen 等CVPR 2021
- Towards Building A Group-based Unsupervised Representation Disentanglement FrameworkTao Yang, Xuanchi Ren, Yuwang Wang, Wenjun Zeng 等ICLR 2022 · 被引用 36 次
