Multimodal Gaussian Mixture Variational Autoencoder with Consistency Regularizations
Yarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang, Tingting Zhao, Yun Liao, Yancui Shi
摘要
Variational autoencoder (VAE)-based frameworks possess a natural advantage in modeling the shared and private information inherent in multimodal data. However, current models focus on improving the quality of shared representations from the reconstruction perspective, lacking explicit mechanisms to model their underlying semantic structure. In this paper, we propose the multimodal Gaussian mixture variational autoencoder with consistency regularizations, which introduces a Gaussian mixture prior over the shared latent space to enhance its semantic structure and encourage the formation of cluster-aware latent representations. To address the cross-modal inconsistency problem under missing modality conditions, we propose a cluster-guided regularization strategy that enforces the cross-modal consistency using the pseudo-category labels from unsupervised clustering. Additionally, we design a self-supervised contrastive regularization strategy to align semantically similar representations across modalities. Extensive experiments on MNIST-SVHN and MNIST-CDCB datasets demonstrate that our method significantly outperforms prior state-of-the-art models in generation, classification, and retrieval tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
- Variational Interaction Information Maximization for Cross-domain DisentanglementHyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung KimNeurIPS 2020 · 被引用 65 次
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo 等ICLR 2022 · 被引用 50 次
- Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering StructuresGehui Xu, Jie Wen, Chengliang Liu, Bing Hu 等AAAI 2024 · 被引用 44 次
- Deep Generative Clustering with Multimodal Diffusion Variational AutoencodersEmanuele Palumbo, Laura Manduchi, Sonia Laguna, Daphné Chopard 等ICLR 2024 · 被引用 21 次
相关 Paper
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
- Gaussian Mixture Variational Autoencoder with Contrastive Learning for Multi-Label ClassificationJunwen Bai, Shufeng Kong, Carla P. GomesICML 2022 · 被引用 48 次
- Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic EmbeddingXu Yan, Jun Yin, Jie WenCVPR 2025
- Relating by Contrasting: A Data-efficient Framework for Multimodal Generative ModelsYuge Shi, Brooks Paige, Philip H. S. Torr, N. SiddharthICLR 2021 · 被引用 42 次
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
