Disentangled Cross-Modal Representation Learning with Enhanced Mutual Supervision
Lu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo, Cheng Liang
摘要
Cross-modal representation learning aims to extract semantically aligned representations from heterogeneous modalities such as images and text. Existing multimodal VAE-based models often suffer from limited capability to align heterogeneous modalities or lack sufficient structural constraints to clearly separate the modality-specific and shared factors. In this work, we propose a novel framework, termed D isentangled C ross-M odal Representation Learning with E nhanced M utual Supervision (DCMEM). Specifically, our model disentangles the common and distinct information across modalities and regularizes the shared representation learned from each modality in a mutually supervised manner. Moreover, we incorporate the information bottleneck principle into our model to ensure that the shared and modality-specific factors encode exclusive yet complementary information. Notably, our model is designed to be trainable on both complete and partial multimodal datasets with a valid Evidence Lower Bound. Extensive experimental results demonstrate significant improvements of our model over existing methods on various tasks including cross-modal generation, clustering and classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- Multi-VAE: Learning Disentangled View-common and View-peculiar Visual Representations for Multi-view ClusteringJie Xu, Yazhou Ren, Huayi Tang, Xiaorong Pu 等ICCV 2021 · 被引用 158 次
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
- Multi-View Information-Bottleneck Representation LearningZhibin Wan, Changqing Zhang, Pengfei Zhu, Qinghua HuAAAI 2021 · 被引用 116 次
- Multimodal Generative Learning Utilizing Jensen-Shannon-DivergenceThomas M. Sutter, Imant Daunhawer, Julia E. VogtNeurIPS 2020 · 被引用 105 次
相关 Paper
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva 等ICML 2026
- Learning Multimodal VAEs through Mutual SupervisionTom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth 等ICLR 2022 · 被引用 27 次
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 被引用 6 次
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu 等ACM MM 2020 · 被引用 63 次
- Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal RetrievalJialin Tian, Kai Wang, Xing Xu, Zuo Cao 等SIGIR 2022 · 被引用 19 次
