Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal Retrieval
Jialin Tian, Kai Wang, Xing Xu, Zuo Cao, Fumin Shen, Heng Tao Shen
摘要
Zero-Shot Cross-Modal Retrieval (ZS-CMR) has recently drawn increasing attention as it focuses on a practical retrieval scenario, i.e., the multimodal test set consists of unseen classes that are disjoint with seen classes in the training set. The recently proposed methods typically adopt the generative model as the main framework to learn a joint latent embedding space to alleviate the modality gap. Generally, these methods largely rely on auxiliary semantic embeddings for knowledge transfer across classes and unconsciously neglect the effect of the data reconstruction manner in the adopted generative model. To address this issue, we propose a novel ZS-CMR model termed Multimodal Disentanglement Variational AutoEncoders (MDVAE), which consists of two coupled disentanglement variational autoencoders (DVAEs) and a fusion-exchange VAE (FVAE). Specifically, DVAE is developed to disentangle the original representations of each modality into modality-invariant and modality-specific features. FVAE is designed to fuse and exchange information of multimodal data by the reconstruction and alignment process without pre-extracted semantic embeddings. Moreover, an advanced counter-intuitive cross-reconstruction scheme is further proposed to enhance the informativeness and generalizability of the modality-invariant features for more effective knowledge transfer. The comprehensive experiments on four image-text retrieval and two image-sketch retrieval datasets consistently demonstrate that our method establishes the new state-of-the-art performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation LearningChengxuan Qian, Shuo Xing, Li Li, Yue Zhao 等ICLR 2026 · 被引用 42 次
- CFIR: Fast and Effective Long-Text To Image Retrieval for Large CorporaZijun Long, Xuri Ge, Richard McCreadie, Joemon M. JoseSIGIR 2024 · 被引用 10 次
- Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal ClusteringZhengzhong Zhu, Pei Zhou, Weihong Du, Shiquan Min 等AAAI 2026
相关 Paper
- Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal RetrievalKaiyi Lin, Xing Xu, Lianli Gao, Zheng Wang 等AAAI 2020 · 被引用 50 次
- Correlated Features Synthesis and Alignment for Zero-shot Cross-modal RetrievalXing Xu, Kaiyi Lin, Huimin Lu, Lianli Gao 等SIGIR 2020 · 被引用 22 次
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu 等ACM MM 2020 · 被引用 63 次
- Variational Interaction Information Maximization for Cross-domain DisentanglementHyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, Kee-Eung KimNeurIPS 2020 · 被引用 65 次
- Generalized Zero-Shot Learning via Disentangled RepresentationXiangyu Li, Zhe Xu, Kun Wei, Cheng DengAAAI 2021 · 被引用 88 次
